What does the cache do to tokens per second?
Everything so far has been about bytes. This part is about the number a business actually buys: tokens per second, and the price of each one.
The connection is short. A decode step’s time is its bytes over the bandwidth. The cache is most of those bytes at any interesting context length. So the cache sets the token rate, and it sets it twice over, once through how many sequences fit and once through how long each step takes.
What makes this a good interview topic is that the two obvious goals point in opposite directions. Faster for one user and more users per box are the same knob turned different ways, and a candidate who does not say which one they are optimising has not answered the question.
The 30-second version
First ask what the batch is, because the answer is completely different at batch 1 and at batch 200. Per-user speed is one over the step time, and the step time is the weights plus the batch's cache over the aggregate bandwidth, so at a large batch the fastest fix is to run a smaller one, which costs total throughput and is a business decision, not an engineering one. To get both, you have to move the curve rather than slide along it: fewer bytes per token with an FP8 cache, fewer bytes read per step with sparse attention, or more tokens per read with speculative decoding, which is the one lever that raises per-user speed without touching the batch. Speculative decoding is also the one with a caveat: it costs several times the arithmetic per step, so it pays handsomely at small batch and stops paying once the step is compute bound. And check the obvious things first: whether the context is longer than it needs to be, whether prefix caching is on, and whether prefill is stalling the decodes behind it.What is a serving stack actually optimising?
Four numbers, and it is worth being precise because interviewers use them as shibboleths.
Time to first token is queueing plus prefill: what a user waits before anything appears. Time per output token, or inter-token latency, is the gap between tokens after that, which is one decode step. Throughput is total tokens a second from the hardware. Goodput, the term DistServe introduced, is the request rate you can serve while meeting targets on both latency numbers at once, and it is the only one of the four that captures the trade.
For a sense of what targets look like in practice, NVIDIA’s interactive scenario for Llama 2 70B in MLPerf uses 450 milliseconds for time to first token and 40 milliseconds per output token, which is 25 tokens a second per user, roughly reading speed.
The first step of the figure is a scheduler timeline: three users, chunked prefill interleaved with decode steps, and both metrics marked on it.
Why does adding sequences stop helping?
Sweep the batch and two curves move in opposite directions.
Total tokens a second rises at first, because the weights are read once for the whole batch, and a second sequence gets its arithmetic almost free. Then it flattens, because every extra sequence brings its own cache and reads all of it.
The flattening has a limit that has nothing to do with the model’s arithmetic:
\[\text{tokens/s} \;\to\; \frac{\text{aggregate bandwidth}}{t \cdot \text{bytes per token}}\]At 4k of context on a node of eight H100s that is about 20,000 tokens a second no matter how large the batch gets. And the sweep stops before it: the cache pool from Part 2 runs out at about 320 sequences, where the node is doing roughly 15,000 tokens a second and each user is getting 47.
Per user the rate falls the whole way, from about 190 tokens a second at batch 1 down to those 47. Both of those numbers come out of the same division, which is the next question.
How far does context move the wall?
Both limits move together, because both are the same quantity divided by the same thing.
The pool holds about 1.3 million cached tokens, so the batch is that over the context. The bandwidth ceiling is bandwidth over one sequence’s cache, which is also the context in the denominator. Sixteen times the context is sixteen times fewer sequences and sixteen times fewer tokens a second out of the node.
At 4k this node serves around 320 sequences and on the order of 15,000 tokens a second. At 128k it serves ten, and a few hundred. Nothing about the model changed.
That is the whole reason long-context requests are priced differently, and the reason the compression techniques in Part 3 are worth their complexity rather than being an optimisation you get to later.
What shape is the trade-off?
Plot the two throughputs against each other with the batch as the parameter and something clean falls out: in a purely memory-bound model the frontier is a straight line.
It has to be. Both axes are one over the same step time, so the parametric curve is linear, and its two intercepts are exactly the numbers from two questions ago. Where per-user speed goes to zero, total throughput is the bandwidth ceiling. Where total throughput goes to zero, per-user speed is what one lonely user gets: bandwidth over the weights.
A real server’s curve bends below that line, because of kernel overheads, tensor-parallel all-reduces, scheduling gaps and the arithmetic term. But the skeleton is right, and it makes the trade-off obvious: a latency target is a horizontal line across the chart, and where it meets the curve is the largest batch you are allowed to run.
That is capacity planning for a decode pool, in one sentence. Everything else is trying to move the line.
How do you get more than one token per read?
If a decode step is memory bound, the arithmetic is nearly free, and the obvious trade is to spend some of it.
Speculative decoding does exactly that. A small draft model proposes $k$ tokens, the large model verifies all of them in a single forward pass, and a modified rejection-sampling rule accepts a prefix of them while leaving the output distribution exactly what the large model would have produced alone.
Leviathan and colleagues published it in late 2022, reporting two to three times on T5-XXL, and Chen and colleagues published the same idea independently a few months later, reporting two to two and a half times on Chinchilla 70B.
The accounting is what makes it obviously right for this regime. The verify pass reads the weights once and the cache once, exactly as an ordinary step does. It does $k+1$ times the arithmetic. And it produces the accepted run, which for a draft of five at eighty percent acceptance averages 3.7 tokens.
Unchanged bytes, more tokens: on a memory-bound step that is a multiplier straight through to per-user speed. Medusa’s paper makes the argument in exactly those terms, that the fix for a memory-bandwidth-bound decoder is to raise its arithmetic intensity.
What does the state of the art look like?
The drafts got better, which is where nearly all the gains since have come from.
Medusa attached extra decoding heads to the model itself and reported 2.3 to 2.8 times. EAGLE drafted at the feature level rather than the token level. EAGLE-3 changed the target again, dropping feature prediction for direct token prediction on fused multi-layer features, and reports up to 6.5 times on single-stream decoding, about 1.4 times better than EAGLE-2.
SGLang’s own published benchmark is a useful sanity check on what that means end to end: 158 tokens a second without speculation, 244 with EAGLE-2, 373 with EAGLE-3. It is the recommended setting in their documentation, and both vLLM and TensorRT-LLM ship EAGLE-3 configurations.
The 2026 direction is parallel drafting. P-EAGLE replaces the drafter’s own autoregressive loop with a single parallel pass through a learned shared hidden state, and reports a further speedup on top of EAGLE-3 on every model it was measured on.
When does speculative decoding stop paying?
This is the follow-up question, and it separates people who have read the papers from people who have run the thing.
The gain is the accepted run only while the step stays memory bound. Push the batch up and the verify pass, with its $k+1$ times the query positions, becomes compute bound, and from there the speedup decays toward the ratio of the memory time to the compute time. Add the drafter’s own cost and it can fall below one.
Where that happens depends on the context, and the direction surprises people. At long context the cache is enormous, the pool caps the batch before the arithmetic ever dominates, and speculation keeps paying all the way to the largest batch you can run. At short context the cache is small, the batch can be huge, and speculation loses. MagicDec makes exactly this argument and shows the gains surviving at moderate to large batch for long sequences.
So the interview answer is not “it stops helping at large batch”. It is “it stops helping when the step stops being memory bound, and whether the batch can get there depends on how much cache each sequence is carrying”.
What actually moves the curve?
Put the whole series on one line. A decode step is bytes over bandwidth, so name the term, the factor and the price.
Grouped-query attention divides the bytes per token by eight and is already in the model you are serving, so there is nothing left to collect. An FP8 cache halves them again, for a point or two on hard evaluations. Latent attention is another factor of three or four, but only if you get to choose the model.
Sparse reading divides the bytes read per step, by sixty-four at 128k and by nothing much at 1k, and leaves capacity exactly where it was. Paged blocks and prefix reuse recover the two to four times that fragmentation was eating.
Speculative decoding multiplies tokens per read at small batch and does nothing at large. And a bigger node buys pool and bandwidth linearly, in exchange for money and communication.
The two questions worth asking before any of it: which term is the bottleneck right now, the pool, the bandwidth or the arithmetic, and is the goal one user’s latency or the node’s throughput. Those two questions are most of what a good answer sounds like.
Rapid fire: can you do these from memory?
- Define time to first token, time per output token, throughput and goodput, and say which one captures the trade-off.
- Write the bandwidth ceiling on total tokens per second and evaluate it at 4k and 128k on a node of eight H100s.
- Why is the latency and throughput frontier a straight line in a memory-bound model, and what are its two intercepts?
- Why does speculative decoding preserve the target model's output distribution exactly?
- Give the expected accepted run for a draft of length k at acceptance rate alpha.
- Under what conditions does speculative decoding stop paying, and why does long context delay that point?
- A user complains their tokens arrive slowly while your dashboard shows record throughput. What happened, and what do you change?
- Which single technique in this series improves capacity and bandwidth and latency at once, and what does it cost?
Part 6 is the question set: three levels of questions on everything above, with answer sketches rather than essays.