Fundamentals / September 23, 2026 / 7 min read

Why Local LLMs Slow Down as Conversations Grow

The KV cache stops your GPU from recalculating the past, but it won't stop it from reading it.

On this page
  1. 1. Reading Prompts vs. Writing Words: Why Parallelism Dies
  2. 2. What the KV Cache Saves (and What It Doesn’t)
  3. 3. The Theoretical Speed Ceiling
  4. 4. Measured Context Slowdown
  5. 5. How Frontier Architectures Fight Back
  6. 6. Practical Tuning for Local Runtimes
  7. The Takeaway

On turn one of a chat, a 27B model in llama.cpp flies at 60 tokens per second. Fifty turns into a coding session, it crawls at 22. You check nvidia-smi—GPU utilization is high, yet 6 GB of VRAM remains completely free.

The common explanation is that the model has “more text to calculate.” But modern GPUs have hundreds of TFLOPs of compute; running dot-products against past words takes microseconds. The real bottleneck is physical: memory bandwidth saturation. Generating words one by one turns your high-speed GPU into a memory-bound traffic jam.


1. Reading Prompts vs. Writing Words: Why Parallelism Dies

To understand why GPUs slow down as conversations grow, you have to separate prompt ingestion from token generation.

  • Prompt Ingestion (Prefill) works like a conveyor belt. When you hit enter, the GPU digests your entire prompt at once. Because hundreds or thousands of prompt words enter the model together, the GPU loads its weights from memory once and processes all of them in parallel. Your GPU’s compute cores stay fully saturated and working at peak efficiency.
  • Token Generation (Decode) works like a courier running back and forth across a narrow bridge. To write a response, the model must produce words sequentially—each new word depends on the word that came right before it. To emit a single word, the GPU cannot batch. It must stream every gigabyte of model weights and every megabyte of chat history out of VRAM and through the memory bus—just to predict one word. The cores finish the math in microseconds and spend the rest of the cycle sitting idle, waiting on memory bandwidth.

2. What the KV Cache Saves (and What It Doesn’t)

To predict the next word, a transformer checks how its current position relates to every earlier word in the conversation. Without caching, generating word 1,000 would require recalculating all 999 earlier words from scratch. Generating word 1,001 would recalculate 1,000 words. That compounds into a massive waste of compute.

The Key-Value (KV) cache fixes this by saving the computed state (the Key and Value vectors) of every past word directly into VRAM the first time it is processed.

Watch the KV cache pull ahead

Local models keep data private

Start both together. Watch caching skip repeated work and finish first.

BuildReuseIllustrative timing

Without cache

0.0 s · 0/12 tokens

Ready

With cache

0.0 s · 0/12 tokens

Ready
Details
Cache memory0 B
Total K/V work saved0 positions

K/V means keys and values: the attention data computed for each token. Words stand in for token positions here; a model’s actual tokenizer may split them differently.

0 cached positions

Representative positions are shown for large contexts.

Both lanes use the same teaching weights: 60 ms to build a position’s K/V, 15 ms to read a position, and 150 ms of other work per output token. Only the cached lane skips rebuilding old K/V. Above 16 context positions, both per-position costs are scaled by 16/context to keep the race watchable. These illustrative times are not hardware measurements or predicted speedups.

Attention still reads the context in both lanes. Step advances the shared clock to the next token arrival; it does not force the two outputs to stay in sync. The last sampled token remains uncached because no further decode step consumes it.

Cache bytes = positions × 32 layers × 2 (K + V) × 8 KV heads × 128 head dimension × bytes per element. Batch size 1; no sharding, quantization scales, or metadata overhead.

As the simulation shows, caching skips repeated math and keeps latency viable. But it introduces an unavoidable memory trade-off:

The KV cache stops your GPU from recalculating the past. It does not stop it from reading the past.

Every new word requires:

  1. Calculating the new word’s representations.
  2. Appending them to the cache in VRAM.
  3. Streaming all previously cached words across the memory bus to calculate attention.

Model weights remain static (~15–30 GB for typical local models), but the KV cache expands with every generated word. Every turn added to history increases the volume of data the GPU must pull across its memory bus for every single subsequent word.


3. The Theoretical Speed Ceiling

You can estimate your hardware’s maximum possible generation speed using a single back-of-the-envelope formula:

Max Generation Speed (words/sec) ≈ Memory Bandwidth (GB/s) / (Model Size (GB) + Cache Size (GB))

On a popular workstation GPU like the RTX 3090 (24 GB VRAM, 936 GB/s memory bandwidth) running a 14 GB model:

  • Turn 1 (Empty cache):
    936 GB/s / (14 GB + 0 GB) ≈ 66.8 tokens/sec
  • Turn 50 (64K context, ~8 GB standard cache):
    936 GB/s / (14 GB + 8 GB) ≈ 42.5 tokens/sec

Bandwidth saturation alone cuts generation throughput by ~36%, before accounting for software overhead, kernel launch delays, or attention calculations.

Capacity vs. Bandwidth

Never confuse memory capacity with memory bandwidth:

  • Capacity (GB): Does the model and its cache fit in VRAM without crashing (OOM)?
  • Bandwidth (GB/s): How quickly can the GPU read that data on every forward pass?

Having 6 GB of free VRAM means your card will not crash with an Out-of-Memory error; it does not make the memory bus any wider.

The PCIe Spillover Cliff

What happens when your conversation outgrows physical VRAM? Runtimes like llama.cpp spill excess cache into system RAM over your motherboard’s PCIe slot. But look at the bandwidth differences:

  • Onboard GDDR6X VRAM: ~936 GB/s
  • System DDR5 RAM: 50–90 GB/s
  • PCIe 4.0 x16 Slot: ~31.5 GB/s

The moment even 500 MB of cache spills across PCIe, generation speed falls off a cliff—dropping from 40 tokens/sec to sub-2 tokens/sec.


4. Measured Context Slowdown

Here is what this slowdown looks like on real silicon, measured on a dual-RTX 3090 workstation running Qwen3.8-27B Q4_K_M in llama.cpp (with FlashAttention and a Q8_0 KV cache):

Qwen3.8 27B Q4_K_M · Standard (no MTP) — existing KV-cache depth; 1K = 1,024 tokens
KV depthPrompt ProcessingToken Generation
01,550.49 t/s58.45 t/s
2,0481,501.02 t/s57.96 t/s
4,0961,487.12 t/s57.31 t/s
8,1921,441.90 t/s55.69 t/s
16,3841,359.20 t/s53.57 t/s
32,7681,233.63 t/s49.66 t/s
65,5361,032.98 t/s42.83 t/s
131,072774.93 t/s33.45 t/s
262,144512.93 t/s22.68 t/s

Measured Token Generation shows the steady drop as history deepens:

  • 0K Context: 58.45 tokens/sec
  • 8K Context: 55.69 tokens/sec (–4.7%)
  • 32K Context: 49.66 tokens/sec (–15.0%)
  • 64K Context: 42.83 tokens/sec (–26.7%)
  • 131K Context: 33.45 tokens/sec (–42.8%)
  • 262K Context: 22.68 tokens/sec (–61.2%)

At 256K context, generation speed drops to less than 40% of its initial rate.

Prompt processing (prefill) also falls—from 1,550 down to 513 tokens/sec—because each new prompt token has to attend across hundreds of thousands of prior words.

Yet Qwen3.8 holds up far better than pure standard transformers like Llama 3 or Gemma. That resilience comes from how modern architectures handle context memory.


5. How Frontier Architectures Fight Back

Modern architectures tackle memory bandwidth saturation through four main strategies:

  1. Compressing the Cache (DeepSeek’s MLA): Multi-Head Latent Attention avoids storing full key and value vectors. Instead, it compresses them into a compact latent representation in memory, then reconstructs them on-chip during generation. This slashes per-token memory traffic by 70–80%.
  2. Reading Only Relevant Chunks (MiniMax M3): Uses a lightweight indexer to look at history in blocks (128 words each) and only loads the top-16 most relevant blocks into attention.
    The catch: Reading fewer blocks keeps generation fast, but the full conversation must still reside in VRAM (~12.3 GB at 100K) so the indexer can search it. It fixes generation bandwidth, not memory capacity.
  3. Fixed-Size Rolling Memory (Linear Recurrence / DeltaNet): Instead of keeping a list of every token, recurrent layers maintain a single fixed-size state matrix that updates as the conversation proceeds:
    State(t) = α(t) × State(t - 1) + β(t) × (k(t)ᵀ × v(t))
    Because the summary size is fixed, memory never grows with conversation length (O(1) memory). Turn 100,000 uses the exact same memory and bandwidth as Turn 1.
  4. Hybrid Stacking (Qwen3.8, DeepSeek V4.1): Stacks recurrent layers (handling bulk context at zero cache growth) with standard attention layers (for exact needle-in-a-haystack recall). Qwen3.8 pairs 48 DeltaNet layers with 16 attention layers; only those 16 layers grow a KV cache.
Architecture / Model Cache Growth (Bytes/Token) 100K Context Cache Size Primary Mechanism Bandwidth Bottleneck Solved? VRAM Capacity Solved?
Standard Dense (Llama 3.3 70B) ~131,072 B (128 KiB) 13.1 GB Full attention across all layers (BF16) ❌ No (O(N) growth) ❌ No
Sparse Reads (MiniMax M3) ~122,880 B (120 KiB) 12.3 GB Top-16 sparse block retrieval ✅ Yes (reads top blocks) ❌ No (stores full history)
Hybrid Recurrent (Qwen3.8-27B) 65,536 B (64 KiB) 6.55 GB 48 DeltaNet + 16 attention layers ✅ Yes (67% fewer attention reads) ✅ Yes (48 layers take 0 B cache)
Compressed MLA (Kimi K3) 27,648 B (27 KiB) 2.76 GB 69 recurrent + 24 compressed MLA layers ✅ Yes (Recurrence + MLA latents) ✅ Yes
Stacked CSA2 (DeepSeek V4.1 Flash) 890 B 0.089 GB (89 MB) Sparse indexer + FP4 native cache ✅ Yes (Under 1 GB at 1M context) ✅ Yes

6. Practical Tuning for Local Runtimes

You do not have to accept a 50% context slowdown as fixed. Here are four practical levers to protect generation speed in llama.cpp and vLLM:

1. Quantize the KV Cache (-ctk and -ctv)

Quantizing model weights (e.g. Q4_K_M) does not compress the KV cache. By default, runtimes store cache entries at uncompressed 16-bit precision (FP16).

Pass explicit cache quantization flags:

# Halve cache memory bandwidth with negligible quality loss
llama-server -m model.gguf -c 65536 -ctk q8_0 -ctv q8_0

# Or cut cache bandwidth by 75% for extreme sequence lengths
llama-server -m model.gguf -c 131072 -ctk q4_0 -ctv q4_0
  • q8_0 KV Cache: Slashes cache memory traffic by 50% with virtually zero perplexity impact. Free performance past 16K context.
  • q4_0 KV Cache: Slashes cache bandwidth by 75%. Expect a minor accuracy penalty on needle retrieval, but it restores 30–40% of generation speed past 64K context.

2. Force FlashAttention (-fa)

Always include --flash-attn (-fa). FlashAttention keeps intermediate attention math inside fast on-chip SRAM (~19 TB/s), avoiding unnecessary VRAM read/write round-trips as sequence length grows.

3. Use Speculative Decoding & Multi-Token Prediction (MTP)

Native MTP (built into Qwen3.8) and speculative decoding evaluate candidate words in a single parallel verification step. Emitting 2–3 accepted words per memory transfer cycle amortizes static weight and KV cache reads, sustaining high effective throughput even deep in history.

4. Choose Hybrid Models for Long Sessions

For agentic coding, multi-turn tool loops, or repository analysis, avoid pure dense models. A hybrid model like Qwen3.8-27B retains 75% of its layers in zero-growth recurrent state, keeping cache sizes compact and generation responsive across long sessions.


The Takeaway

When your local model slows down, it is neither out of compute nor out of VRAM.

Every generated word forces the GPU to stream static weights and growing history across the memory bus. As context expands, memory bandwidth saturates, starving compute cores and slowing output.

To keep long conversations fast:

  1. Compress the cache: Use -ctk q8_0 -ctv q8_0 to halve cache bandwidth.
  2. Keep data on-chip: Turn on FlashAttention (-fa).
  3. Pick the right architecture: Choose hybrid recurrent models (Gated DeltaNet / KDA) over pure dense transformers for deep-context tasks.

Further Reading & Benchmarks

About the author

Harris Oldroyd

Independent self-taught builder and researcher

I learn systems from first principles, build them, and measure them before writing about them. The notebook covers local AI hardware and inference, systematic trading research, and the software that keeps both repeatable.