How much memory does the cache actually take?
This is the arithmetic an interviewer will make you do on a whiteboard, and it is the one place where being approximately right is not good enough. The formula has five factors and every one of them is a design decision somebody made.
The good news is that once you can write it, a whole family of questions becomes arithmetic. How many users fit on this box. Whether 128k context is affordable. Why the second GPU did not double the batch. What FP8 buys.
Everything below is Llama 3 70B on H100s, the same shape as Part 1: 80 layers, 8 key-value heads, head dimension 128, 70.55 billion parameters.
The 30-second version
Start with bytes per token: two, for key and value, times layers, times key-value heads, times head dimension, times bytes per number. For Llama 3 70B in BF16 that is 320 KiB. The node has 640 GB; the BF16 weights take 141 GB, the runtime holds back roughly a tenth for activations, graphs and workspace, and what is left, a little over 400 GB, is the cache pool. Divide: about 1.3 million cached tokens in total, which is around 40 sequences at 32k, 320 at 4k, and 10 at 128k. Then check the second constraint, because capacity is only half of it: every decode step reads the whole pool, so the step time is those bytes over the aggregate bandwidth, and that is what caps tokens per second no matter how many sequences fit.Where does the per-token number come from?
Every token stores a key and a value, for every layer and every key-value head, each of them $d_h$ numbers wide:
\[\text{bytes per token} = 2 \cdot L \cdot n_{kv} \cdot d_h \cdot s\]The 2 is key and value. $L$ and $n_{kv}$ and $d_h$ are the model’s shape. $s$ is bytes per number: 2 for BF16, 1 for FP8.
For this model: $2 \times 80 \times 8 \times 128 \times 2 = 327{,}680$ bytes, 320 KiB. Set $n_{kv}$ back to 64, the way attention was written in 2017, and it is 2.5 MiB.
Two things are missing from that formula and their absence matters. The batch is not in it, because it multiplies later. The 64 query heads are not in it, because queries are never stored.
The figure below is the whole of this part: sliders for the shape, then per sequence, then per node, then the map of what fits.
When does the cache overtake the weights?
Multiply by the context and you have one sequence. Multiply by the batch and you have the server. Both numbers are chosen by users, not by you: how long their conversation is, and how many of them arrive at once.
The weights are 141 GB and they do not move. The cache equals them when
\[b \cdot t = \frac{2P}{2 L n_{kv} d_h s} = 430{,}000 \text{ tokens}\]for this model. One sequence would have to reach 430k tokens to get there on its own. A batch of 128 gets there at 3,364 tokens each, which is a medium-length chat.
That is the sentence to carry out of here: on a busy server the cache is not a rounding error next to the weights, it is the larger of the two, and it got that way at ordinary conversation lengths.
What fits on one node?
A single H100 cannot hold this model at all. 141 GB of weights against 80 GB of HBM, so the smallest sensible deployment is a node of eight with tensor parallelism, each GPU holding an eighth of every weight matrix and one of the eight key-value heads.
Then the accounting per GPU. vLLM takes 92% of the card by default, which is the gpu_memory_utilization setting. The weight shard is 17.6 GB. A couple of gigabytes go to activations, CUDA graphs and kernel workspace.
What is left is the cache pool, and across the node it is a little over 400 GB. In tokens, about 1.3 million of them, and that single number answers most capacity questions.
Divide by the context: about 320 sequences at 4k, 40 at 32k, 10 at 128k. Ten. That is the whole node, running one of the most-deployed open models in the world, serving ten people at its advertised context length.
Why is capacity only half the problem?
Because those bytes are not stored and forgotten. Every decode step reads the whole cache back, for every sequence in the batch, to compute one row of attention each.
So the step’s traffic is the weights plus the cache:
\[\text{bytes per step} = 2P + b \, t \cdot 2 L n_{kv} d_h s\]At 4k and batch 32 that is 184 GB, of which the cache is 23%. At 4k and batch 321, the pool’s limit, it is 572 GB and the cache is three quarters.
Batching amortises the weights across sequences. It never amortises the cache, because each sequence reads only its own. That asymmetry is the reason the next question has the answer it does.
Why can decode never reach the roof?
Part 2 of the GPU series drew the roofline and found decode stuck far down the memory slope. Here is the same fact from the cache’s side.
Arithmetic intensity is FLOPs per byte. Add sequences and the FLOPs grow, but so do the cache bytes, at exactly the same rate. So intensity rises with the batch to a ceiling and stops:
\[I \to \frac{2P + 4 L n_h d_h t}{t \cdot 320\,\text{KiB}} \quad \text{as } b \to \infty\]That ceiling is about 60 FLOPs per byte at 8k of context and 11 at 128k, exactly the numbers the GPU series quoted. The H100’s ridge is 295.
No batch size reaches the roof, and a longer context pushes you further from it. Decode is memory bound permanently, and that is a property of the cache, not of the kernel.
What is the token rate that falls out of this?
Put the two halves together. A step moves the weights plus the batch’s cache; the node has 26.8 TB/s of aggregate HBM bandwidth; the step time is one divided by the other.
At 4k and the pool’s limit of about 320 sequences, that is roughly 21 ms per step: on the order of 15,000 tokens a second from the node, and about 47 a second for any one user. At 128k and its limit of ten sequences, the same node yields a few hundred tokens a second in total.
Same model, same hardware, same code. Two orders of magnitude of throughput, decided by one number in the request.
How honest are these numbers?
They are a bandwidth model, and a real server will not hit them. Kernel launch overheads, the all-reduces that tensor parallelism needs at every layer, scheduling gaps, and the fact that nobody runs at exactly the pool’s limit all take their cut. Half of the modelled figure is a reasonable expectation for a tuned stack, and I would say so rather than quote the model as a measurement.
For a calibration point that is measured: NVIDIA’s MLPerf Inference v5.0 submission served Llama 2 70B at 33,072 tokens a second on eight H200s in the server scenario, on much shorter sequences than 4k and with a quantised model. That is what the same shape of arithmetic looks like after a vendor has tuned every layer of it.
What the model is good for is ratios. Double the context and halve the batch, twice the bytes per token and half of everything: those relationships survive every inefficiency, and they are what an interviewer is testing.
Which numbers are worth memorising?
Four, and they are enough to reconstruct the rest.
320 KiB per token for a Llama-70B-shaped model in BF16. 141 GB for its BF16 weights. About 1.3 million cached tokens in a node of eight H100s after the weights and the runtime take their share. And 295 FLOPs per byte for the H100’s ridge, which decode never approaches.
From those you can derive the batch at any context, the step time at any batch, the crossover with the weights, and the intensity ceiling. Everything else in this series is a way of changing one of them.
Rapid fire: can you do these from memory?
- Write the bytes-per-token formula and say why neither the batch nor the query head count is in it.
- At what batch times context does the cache of Llama 3 70B weigh as much as its weights?
- Derive the cache pool of a node of eight H100s running this model, naming everything you subtract.
- How many sequences fit at 4k, 32k and 128k, and why is it the same division three times?
- Why does batching raise arithmetic intensity to a ceiling instead of without limit?
- Give the intensity ceiling at 8k and at 128k, and compare both to the H100's ridge.
- A colleague doubles the batch and sees no more tokens per second. What do you check first?
- Which of these numbers would you refuse to quote as a measurement, and what would you say instead?
Part 3 attacks the formula factor by factor: fewer key-value heads, a compressed latent, fewer bits, a bounded window, and the 2026 answer of keeping every byte but reading almost none of them.