Fundamentals / September 23, 2026 / 7 min read
Why Local LLMs Slow Down as Conversations Grow
The KV cache stops your GPU from recalculating the past, but it won't stop it from reading it.
On this page
On turn one of a chat, a 27B model in llama.cpp flies at 60 tokens per second. Fifty turns into a coding session, it crawls at 22. You check nvidia-smi—GPU utilization is high, yet 6 GB of VRAM remains completely free.
The common explanation is that the model has “more text to calculate.” But modern GPUs have hundreds of TFLOPs of compute; running dot-products against past words takes microseconds. The real bottleneck is physical: memory bandwidth saturation. Generating words one by one turns your high-speed GPU into a memory-bound traffic jam.
1. Reading Prompts vs. Writing Words: Why Parallelism Dies
To understand why GPUs slow down as conversations grow, you have to separate prompt ingestion from token generation.
- Prompt Ingestion (Prefill) works like a conveyor belt. When you hit enter, the GPU digests your entire prompt at once. Because hundreds or thousands of prompt words enter the model together, the GPU loads its weights from memory once and processes all of them in parallel. Your GPU’s compute cores stay fully saturated and working at peak efficiency.
- Token Generation (Decode) works like a courier running back and forth across a narrow bridge. To write a response, the model must produce words sequentially—each new word depends on the word that came right before it. To emit a single word, the GPU cannot batch. It must stream every gigabyte of model weights and every megabyte of chat history out of VRAM and through the memory bus—just to predict one word. The cores finish the math in microseconds and spend the rest of the cycle sitting idle, waiting on memory bandwidth.
2. What the KV Cache Saves (and What It Doesn’t)
To predict the next word, a transformer checks how its current position relates to every earlier word in the conversation. Without caching, generating word 1,000 would require recalculating all 999 earlier words from scratch. Generating word 1,001 would recalculate 1,000 words. That compounds into a massive waste of compute.
The Key-Value (KV) cache fixes this by saving the computed state (the Key and Value vectors) of every past word directly into VRAM the first time it is processed.
Watch the KV cache pull ahead
Local models keep data private
Start both together. Watch caching skip repeated work and finish first.
Without cache
0.0 s · 0/12 tokens
With cache
0.0 s · 0/12 tokens
Details
K/V means keys and values: the attention data computed for each token. Words stand in for token positions here; a model’s actual tokenizer may split them differently.
0 cached positions
Representative positions are shown for large contexts.
Both lanes use the same teaching weights: 60 ms to build a position’s K/V, 15 ms to read a position, and 150 ms of other work per output token. Only the cached lane skips rebuilding old K/V. Above 16 context positions, both per-position costs are scaled by 16/context to keep the race watchable. These illustrative times are not hardware measurements or predicted speedups.
Attention still reads the context in both lanes. Step advances the shared clock to the next token arrival; it does not force the two outputs to stay in sync. The last sampled token remains uncached because no further decode step consumes it.
Cache bytes = positions × 32 layers × 2 (K + V) × 8 KV heads × 128 head dimension × bytes per element. Batch size 1; no sharding, quantization scales, or metadata overhead.
As the simulation shows, caching skips repeated math and keeps latency viable. But it introduces an unavoidable memory trade-off:
The KV cache stops your GPU from recalculating the past. It does not stop it from reading the past.
Every new word requires:
- Calculating the new word’s representations.
- Appending them to the cache in VRAM.
- Streaming all previously cached words across the memory bus to calculate attention.
Model weights remain static (~15–30 GB for typical local models), but the KV cache expands with every generated word. Every turn added to history increases the volume of data the GPU must pull across its memory bus for every single subsequent word.
3. The Theoretical Speed Ceiling
You can estimate your hardware’s maximum possible generation speed using a single back-of-the-envelope formula:
Max Generation Speed (words/sec) ≈ Memory Bandwidth (GB/s) / (Model Size (GB) + Cache Size (GB))
On a popular workstation GPU like the RTX 3090 (24 GB VRAM, 936 GB/s memory bandwidth) running a 14 GB model:
- Turn 1 (Empty cache):
936 GB/s / (14 GB + 0 GB) ≈ 66.8 tokens/sec - Turn 50 (64K context, ~8 GB standard cache):
936 GB/s / (14 GB + 8 GB) ≈ 42.5 tokens/sec
Bandwidth saturation alone cuts generation throughput by ~36%, before accounting for software overhead, kernel launch delays, or attention calculations.
Capacity vs. Bandwidth
Never confuse memory capacity with memory bandwidth:
- Capacity (GB): Does the model and its cache fit in VRAM without crashing (OOM)?
- Bandwidth (GB/s): How quickly can the GPU read that data on every forward pass?
Having 6 GB of free VRAM means your card will not crash with an Out-of-Memory error; it does not make the memory bus any wider.
The PCIe Spillover Cliff
What happens when your conversation outgrows physical VRAM? Runtimes like llama.cpp spill excess cache into system RAM over your motherboard’s PCIe slot. But look at the bandwidth differences:
- Onboard GDDR6X VRAM: ~936 GB/s
- System DDR5 RAM: 50–90 GB/s
- PCIe 4.0 x16 Slot: ~31.5 GB/s
The moment even 500 MB of cache spills across PCIe, generation speed falls off a cliff—dropping from 40 tokens/sec to sub-2 tokens/sec.
4. Measured Context Slowdown
Here is what this slowdown looks like on real silicon, measured on a dual-RTX 3090 workstation running Qwen3.8-27B Q4_K_M in llama.cpp (with FlashAttention and a Q8_0 KV cache):
| KV depth | Prompt Processing | Token Generation |
|---|---|---|
| 0 | 1,550.49 t/s | 58.45 t/s |
| 2,048 | 1,501.02 t/s | 57.96 t/s |
| 4,096 | 1,487.12 t/s | 57.31 t/s |
| 8,192 | 1,441.90 t/s | 55.69 t/s |
| 16,384 | 1,359.20 t/s | 53.57 t/s |
| 32,768 | 1,233.63 t/s | 49.66 t/s |
| 65,536 | 1,032.98 t/s | 42.83 t/s |
| 131,072 | 774.93 t/s | 33.45 t/s |
| 262,144 | 512.93 t/s | 22.68 t/s |
Measured Token Generation shows the steady drop as history deepens:
- 0K Context: 58.45 tokens/sec
- 8K Context: 55.69 tokens/sec (–4.7%)
- 32K Context: 49.66 tokens/sec (–15.0%)
- 64K Context: 42.83 tokens/sec (–26.7%)
- 131K Context: 33.45 tokens/sec (–42.8%)
- 262K Context: 22.68 tokens/sec (–61.2%)
At 256K context, generation speed drops to less than 40% of its initial rate.
Prompt processing (prefill) also falls—from 1,550 down to 513 tokens/sec—because each new prompt token has to attend across hundreds of thousands of prior words.
Yet Qwen3.8 holds up far better than pure standard transformers like Llama 3 or Gemma. That resilience comes from how modern architectures handle context memory.
5. How Frontier Architectures Fight Back
Modern architectures tackle memory bandwidth saturation through four main strategies:
- Compressing the Cache (DeepSeek’s MLA): Multi-Head Latent Attention avoids storing full key and value vectors. Instead, it compresses them into a compact latent representation in memory, then reconstructs them on-chip during generation. This slashes per-token memory traffic by 70–80%.
- Reading Only Relevant Chunks (MiniMax M3): Uses a lightweight indexer to look at history in blocks (128 words each) and only loads the top-16 most relevant blocks into attention.
The catch: Reading fewer blocks keeps generation fast, but the full conversation must still reside in VRAM (~12.3 GB at 100K) so the indexer can search it. It fixes generation bandwidth, not memory capacity. - Fixed-Size Rolling Memory (Linear Recurrence / DeltaNet): Instead of keeping a list of every token, recurrent layers maintain a single fixed-size state matrix that updates as the conversation proceeds:
Because the summary size is fixed, memory never grows with conversation length (O(1) memory). Turn 100,000 uses the exact same memory and bandwidth as Turn 1.State(t) = α(t) × State(t - 1) + β(t) × (k(t)ᵀ × v(t)) - Hybrid Stacking (Qwen3.8, DeepSeek V4.1): Stacks recurrent layers (handling bulk context at zero cache growth) with standard attention layers (for exact needle-in-a-haystack recall). Qwen3.8 pairs 48 DeltaNet layers with 16 attention layers; only those 16 layers grow a KV cache.
| Architecture / Model | Cache Growth (Bytes/Token) | 100K Context Cache Size | Primary Mechanism | Bandwidth Bottleneck Solved? | VRAM Capacity Solved? |
|---|---|---|---|---|---|
| Standard Dense (Llama 3.3 70B) | ~131,072 B (128 KiB) | 13.1 GB | Full attention across all layers (BF16) | ❌ No (O(N) growth) | ❌ No |
| Sparse Reads (MiniMax M3) | ~122,880 B (120 KiB) | 12.3 GB | Top-16 sparse block retrieval | ✅ Yes (reads top blocks) | ❌ No (stores full history) |
| Hybrid Recurrent (Qwen3.8-27B) | 65,536 B (64 KiB) | 6.55 GB | 48 DeltaNet + 16 attention layers | ✅ Yes (67% fewer attention reads) | ✅ Yes (48 layers take 0 B cache) |
| Compressed MLA (Kimi K3) | 27,648 B (27 KiB) | 2.76 GB | 69 recurrent + 24 compressed MLA layers | ✅ Yes (Recurrence + MLA latents) | ✅ Yes |
| Stacked CSA2 (DeepSeek V4.1 Flash) | 890 B | 0.089 GB (89 MB) | Sparse indexer + FP4 native cache | ✅ Yes (Under 1 GB at 1M context) | ✅ Yes |
6. Practical Tuning for Local Runtimes
You do not have to accept a 50% context slowdown as fixed. Here are four practical levers to protect generation speed in llama.cpp and vLLM:
1. Quantize the KV Cache (-ctk and -ctv)
Quantizing model weights (e.g. Q4_K_M) does not compress the KV cache. By default, runtimes store cache entries at uncompressed 16-bit precision (FP16).
Pass explicit cache quantization flags:
# Halve cache memory bandwidth with negligible quality loss
llama-server -m model.gguf -c 65536 -ctk q8_0 -ctv q8_0
# Or cut cache bandwidth by 75% for extreme sequence lengths
llama-server -m model.gguf -c 131072 -ctk q4_0 -ctv q4_0
q8_0KV Cache: Slashes cache memory traffic by 50% with virtually zero perplexity impact. Free performance past 16K context.q4_0KV Cache: Slashes cache bandwidth by 75%. Expect a minor accuracy penalty on needle retrieval, but it restores 30–40% of generation speed past 64K context.
2. Force FlashAttention (-fa)
Always include --flash-attn (-fa). FlashAttention keeps intermediate attention math inside fast on-chip SRAM (~19 TB/s), avoiding unnecessary VRAM read/write round-trips as sequence length grows.
3. Use Speculative Decoding & Multi-Token Prediction (MTP)
Native MTP (built into Qwen3.8) and speculative decoding evaluate candidate words in a single parallel verification step. Emitting 2–3 accepted words per memory transfer cycle amortizes static weight and KV cache reads, sustaining high effective throughput even deep in history.
4. Choose Hybrid Models for Long Sessions
For agentic coding, multi-turn tool loops, or repository analysis, avoid pure dense models. A hybrid model like Qwen3.8-27B retains 75% of its layers in zero-growth recurrent state, keeping cache sizes compact and generation responsive across long sessions.
The Takeaway
When your local model slows down, it is neither out of compute nor out of VRAM.
Every generated word forces the GPU to stream static weights and growing history across the memory bus. As context expands, memory bandwidth saturates, starving compute cores and slowing output.
To keep long conversations fast:
- Compress the cache: Use
-ctk q8_0 -ctv q8_0to halve cache bandwidth. - Keep data on-chip: Turn on FlashAttention (
-fa). - Pick the right architecture: Choose hybrid recurrent models (Gated DeltaNet / KDA) over pure dense transformers for deep-context tasks.
Further Reading & Benchmarks
- Qwen3.8-27B Architecture & Local Benchmarks — Alibaba’s 3:1 hybrid recurrence model, native MTP, and dual-GPU scaling.
- RTX 3090 for Local AI Inference — Single- and dual-GPU context sweeps, bandwidth limits, and power tuning.
- DeepSeek V4.1 Flash Architecture Analysis — How CSA2, FP4 caching, and sparse routing fit 1M context into under 1 GB VRAM.
- Kimi K3 Explained — Breakdown of Moonshot AI’s 2.8T MoE with KDA recurrent memory.
- Long-Context and Attention Explained (RoPE and YaRN) — How position embeddings distort attention angles as context stretches.
- Local AI Token Speed and Memory Calculator — Calculate exact KV cache footprints and predicted throughput for your hardware.