How do you make the cache smaller?
Part 2 left the formula standing: two, times layers, times key-value heads, times head dimension, times bytes per number. Four factors, and the interesting thing about the last decade of attention research is that somebody has attacked every one of them.
There is also a fifth move that does not appear in the formula at all, and it is where 2026 went: keep every byte and stop reading most of them.
The rule for this part is that every scheme is measured on the same model at the same context, so the numbers are comparable. Llama 3 70B’s shape, one sequence, 128k of context.
The 30-second version
Four families. Share key-value heads across query heads: multi-query keeps one pair per layer and loses measurable quality, grouped-query keeps one per group of eight and loses almost none, which is why every open model ships with it. Compress instead of sharing: multi-head latent attention caches a 512-wide latent plus a 64-wide component carrying the rotation, and folds the up-projections into the query and output projections so the keys and values are never built; it is another factor of three or four on top of grouped-query, but it has to be trained in. Use fewer bits: FP8 halves everything for one flag, and below eight bits the layout starts to matter, because keys have outlier channels and values do not. Bound the cache: a sliding window makes it constant in the context, at the cost of everything outside the window, and it only works if you pin the first few tokens as attention sinks. Then the 2026 answer, which is orthogonal to all of them: sparse attention keeps the whole cache resident and reads only the top few thousand tokens per step, so capacity is unchanged and bandwidth falls by the ratio of context to k.What did the cache look like before grouped-query attention?
In the original layout every query head had a key head and a value head of its own. On this shape that is $2 \times 80 \times 64 \times 128 \times 2$ bytes per token: 2.5 MiB.
At 128k of context, one sequence carries 344 GB of cache. The model it is talking to is 141 GB. Eight such users would need four nodes of H100s, for eight people.
This is the number that made long context a research demonstration rather than a product, and it is worth stating out loud in an interview, because it frames everything that follows as a response to a specific, quantified problem.
What do you actually give up by sharing key-value heads?
Shazeer’s multi-query attention, in 2019, keeps one key head and one value head for the whole layer, whatever the number of query heads. A sixty-four fold cut here, and the quality cost is real but small: on the paper’s WMT14 English to German setup, dev perplexity moves from 1.424 to 1.439 and BLEU from 27.7 to 27.5 at beam 1.
Grouped-query attention, from Ainslie and colleagues in 2023, puts one key-value head in front of a group of query heads. Llama 3 uses eight groups of eight.
Their table is the one to remember, because it has both columns. On T5-XXL, averaged over five summarisation benchmarks, multi-head scores 47.2, multi-query 46.6, and grouped-query with eight heads 47.1. Inference time per sample: 1.51 seconds, 0.24, and 0.28. Grouped-query keeps nearly all of multi-query’s speed and nearly all of multi-head’s quality.
The paper also gave the recipe for getting there from a checkpoint you already have: mean-pool the key and value projections within each group, then uptrain for five percent of the original pretraining compute. It is one of the cheapest architecture changes in the field, and the arithmetic is unchanged; the keys are simply read by more than one query head.
How does a latent cache work?
Multi-head latent attention asks a different question. Instead of keeping fewer heads, keep a smaller thing and rebuild the heads from it.
DeepSeek-V2 projects each token down to a 512-wide latent, caches that, and reconstructs all 128 key and value heads with up-projections when attention runs. Rotary embeddings do not commute with the compression, so a separate 64-wide component carries the rotation, giving 576 cached numbers per layer per token.
The part that makes it free at inference is absorption. Since $q^\top (W^{UK} c) = (q^\top W^{UK}) c$, the up-projection folds into the query projection, and the value up-projection folds into the output projection. The keys and values are never materialised at all; attention runs directly against the latent.
On this figure’s 80-layer shape that is 90 KiB per token against grouped-query’s 320 KiB. DeepSeek’s own claim for V2 is a 93.3% smaller cache than their previous 67B model and 5.76 times the maximum generation throughput, and their framing of the size is exact: 576 numbers per layer is grouped-query attention with 2.25 groups, at a quality they measure above multi-head.
The catch is that it is architectural. You cannot switch it on for a checkpoint that was not trained with it, which is why MLA shows up in DeepSeek and Kimi and not in the Llama you already have.
How few bits can the cache survive?
The last factor is the size of each number, and it is the only one you can change on a model you did not train.
FP8 halves the cache for one flag in every serving engine. vLLM’s own April 2026 write-up on FP8 KV cache is the most useful public evidence: with e4m3 and per-head scales they report recovering roughly 97 to 99 percent of baseline accuracy on long-context evaluations, and losing a point or two on the hardest reasoning benchmarks. That is a good summary of the state of it: a production default with a caveat you should be able to name.
Below eight bits the layout matters. KIVI’s observation is that keys have a few channels whose magnitudes are far larger than the rest, so a per-token scale would spend its whole range on those channels; keys should be quantised per channel. Values have no such structure, and attention is sparse over them, so per-token quantisation confines the error to tokens nobody was attending to anyway.
With that asymmetry they reach two bits, and report 2.35 to 3.47 times the throughput from the larger batches it allows.
KVQuant reports under 0.1 perplexity lost at three bits, with per-channel keys quantised before the rotation is applied. And the error is not spread evenly: the tokens that hurt most are the first few, the ones every query attends to, which is the thread the next question pulls on.
What does a window forget?
A different bargain: bound the cache instead of shrinking it. Sliding-window attention lets each token attend only to the last $W$ positions, so the cache stops growing once it reaches $W$. Mistral 7B shipped with a window of 4,096 and reported an eightfold cache reduction at 32k tokens.
Then the failure mode, which is the interesting part. The softmax has to put its probability mass somewhere even when nothing in the window is relevant, and in a trained model the place it dumps that mass is the first few tokens, which every position can see. Evict them and the model falls apart: StreamingLLM measured a 1,024-token window on Llama-2-13B at 65k tokens reaching a perplexity of 5158, and the same window with four initial tokens pinned reaching 5.40.
Four tokens. That is the whole fix, and with it they run stably to four million tokens of streaming input.
2026 models take the middle road rather than the pure form. Gemma 3 alternates five local layers of 1,024 width per global layer; gpt-oss alternates one for one with a 128-wide window and a learned per-head sink. The cost is not perplexity, it is retrieval: what falls outside every window in every layer is gone, and no prompt can bring it back.
Can you keep everything and read almost none of it?
This is the 2026 answer and the one most candidates have not internalised. Storage and traffic were always two separate bills, and nothing forces you to pay them at the same rate.
DeepSeek Sparse Attention, shipped in DeepSeek-V3.2, keeps the whole cache and reads part of it. A small lightning indexer scores every earlier token against the current query using ReLU-gated dot products, cheaply enough to run in FP8, and the main attention runs over the top $k$ alone. The shipped configuration uses 64 indexer heads and $k = 2048$.
Storage is unchanged, so Part 2’s capacity arithmetic still holds exactly. Bytes read per step fall by the ratio of the context to $k$, which at 128k is sixty-four, and the attention FLOPs fall with them.
Native Sparse Attention, from the same lab earlier in 2025, is the trained-from-scratch cousin: three branches, compressed blocks for the coarse view, selected blocks for the fine one, and a sliding window for locality, summed. It reports 11.6 times faster decoding at 64k tokens.
Note what the indexer costs. It still scores every earlier token, so it is linear in the context rather than constant. Sparse reading makes long context cheap per token; it does not make it free.
How do these compose?
They multiply, because they act on different factors, and a strong answer says so with numbers.
Grouped-query attention divides the bytes by eight and is already in the model you are serving. An FP8 cache halves what is left. Sparse reading divides the bytes read per step by another sixty-four at 128k. Together: 344 GB resident and read per step in the 2017 layout, against 21 GB resident and 336 MB read per step with all three.
The last step of the figure puts all of them on one axis at 128k. Two columns, and they are different columns: what is resident decides how many users fit on the node, and what is read per step decides how fast each of them is served.
Which one would you actually reach for?
The order is not the order they were invented in, it is the order of what you are allowed to change.
If you are serving a checkpoint, you have two levers: the bits, and whether your engine supports a sparse-attention variant the model was trained for. Turn on the FP8 cache, measure your own evaluations rather than trusting a general claim, and stop there unless the numbers force more.
If you are choosing a model, grouped-query attention is table stakes and latent attention is a genuine differentiator at long context. It is fair to ask why a model without it is worth the extra memory.
And if you are training one, everything is on the table, which is exactly why the frontier labs’ architecture choices in 2025 and 2026 read like a list of cache optimisations. That is not a coincidence. It is what happens when the serving bill starts writing the model card.
Rapid fire: can you do these from memory?
- Give the per-token cache for multi-head, multi-query and grouped-query attention on the same 80-layer shape.
- What does grouped-query attention change about the arithmetic of attention? Be careful.
- How do you convert a multi-head checkpoint into a grouped-query one, and what does it cost?
- Why does multi-head latent attention need a separate component for the rotation?
- State the absorption identity and say what it saves.
- Why are keys quantised per channel and values per token?
- Why does a sliding window collapse without attention sinks, and how many tokens fix it?
- Which of these techniques leaves the memory footprint completely unchanged, and what does it change instead?
Part 4 leaves the model alone and looks at the server: how the bytes are packed, shared between users, evicted when the pool fills, and shipped to another GPU entirely.