Can you answer these without the diagrams?
Five parts of derivation are worth one hour of being asked about them out loud, so this part is the self-test. Three levels, and they fail differently.
Level 1 fails when you cannot state the mechanism precisely: what is stored, what is computed, where the bytes go. Level 2 fails when you can recite a technique but cannot say what breaks when you turn its knob. Level 3 has no clean answer at all, and the honest response is a position plus the measurement that would change it.
Every question here has a wrong answer that a half-prepared candidate gives, and the sketches are what a strong answer contains, not an essay.
The 30-second version
Treat the open question as an invitation to structure the topic rather than to recite it. There are three things worth saying and they take about a minute: what it is and why it can exist at all, which is that keys and values depend only on their own token; what it costs, which is two times layers times key-value heads times head dimension bytes per token, resident and re-read on every step; and what everyone does about it, which is fewer heads, fewer bits, a bounded window, or reading less of it. Then stop and let them pick the branch. The candidates who go badly here are the ones who start describing attention from scratch.Level 1: can you state the mechanism?
-
Which tensors go in the cache and which do not, and what property of them decides it?
answer sketch
Keys and values, per layer and per key-value head. Not queries. The reason is dependence, not size: the key and value of token j are functions of token j and its position, fixed the moment that token passed through the layer, so they are identical for every later step that reads them. The query is a function of the token being generated, so there is nothing to reuse. A candidate who says "QKV cache" has never written the loop.
-
Write the bytes per token for a KV cache, then evaluate it for Llama 3 70B in BF16. Why is neither the batch size nor the query head count in the formula?
answer sketch
Two, for key and value, times layers, times key-value heads, times head dimension, times bytes per element. For 80 layers, 8 key-value heads and head dimension 128 in BF16 that is 320 KiB. The batch is absent because it multiplies afterwards, once per sequence. The query heads are absent because queries are not stored, which is exactly why grouped-query attention divides the cache while leaving the model's expressive width alone.
-
During one decode step, exactly what does attention compute for a single sequence, and how does that differ from the same layer during prefill?
answer sketch
One row of the causal matrix: one query per head, scored against every cached key, softmaxed, and used to average the cached values. Prefill computes the entire lower triangle at once for all prompt tokens. The arithmetic per token is nearly identical in the two phases; what differs by orders of magnitude is the bytes per token, because prefill shares one read of the weights across the whole prompt and decode does not.
-
List what a decode step reads from HBM, largest first, for a 70B model at batch 32 and 4k context. Which of those are shared across the batch?
answer sketch
The weights, 141 GB in BF16, read once for the whole batch. Then the cache, 43 GB at that batch and context, read once per sequence and shared with nobody. About 184 GB in total. The distinction is the entire theory of batching: the first term amortises and the second does not, which is why intensity rises to a ceiling and stops.
-
Where does the cache live in the memory hierarchy, and what of it is in shared memory while the attention kernel runs?
answer sketch
All of it in HBM, because it has to survive between kernels. During the kernel, blocks of keys and values stream through shared memory a tile at a time while the query block and the running softmax statistics stay in registers, and nothing of the cache is resident above HBM between steps. That is why the cache appears in the roofline's byte count on every single step and the attention score matrix does not.
-
Give the numbers cached per token, per layer, under multi-head, grouped-query and latent attention on the same model shape. Not the bytes: the count.
answer sketch
Multi-head is two times heads times head dimension, so 64 heads at 128 gives 16,384 per layer. Grouped-query with 8 key-value heads gives 2,048. Latent attention caches one 512-wide latent plus a 64-wide component that carries the rotation, so 576, which the DeepSeek paper points out is grouped-query attention with 2.25 groups. The bytes then follow from the element size.
-
What is a block in paged attention, how large is it, and what is in the block table?
answer sketch
A fixed number of tokens' worth of keys and values for every layer and head, sixteen tokens by default in vLLM. The table maps each sequence's logical block index to a physical block in the pool, plus how many slots of the last block are filled. The kernel gathers through it, so a sequence's cache need never be contiguous, and the waste is bounded by one partly filled block per sequence.
-
Why is there no KV cache during training?
answer sketch
Training processes a whole sequence in one pass with teacher forcing, so every key and value is computed once and consumed immediately by the same forward pass. There is no second step that wants them again. What training stores instead is activations for the backward pass, which is a larger and differently shaped problem, and the quadratic part of it is what FlashAttention removed.
Level 2: can you reason about the trade-offs?
-
Doubling the batch from 4 to 8 nearly doubles your throughput. Doubling it from 256 to 512 does almost nothing. Explain both, in the same sentence if you can.
answer sketch
At batch 4 the step's bytes are almost entirely the weights, which do not grow with the batch, so twice the sequences is twice the tokens for the same read. By batch 256 the cache dominates the read and grows exactly in proportion to the batch, so the step time grows with it and the token rate flattens. The crossover is where the batch's cache equals the weights, and it moves left as the context grows.
-
You switch the cache to FP8. What improves, and by how much, on capacity and on per-user latency? Do not say "everything halves".
answer sketch
Capacity does halve the bytes, so the batch at a given context doubles, or the context at a given batch does. Latency improves by much less, because the step also reads the weights: only the cache share of the step's bytes is halved. At a small batch and short context that is a few percent; at the pool's limit, where the cache is most of the read, it approaches a factor of two. The cost is accuracy, concentrated on the hardest evaluations rather than spread evenly.
-
Grouped-query attention with eight groups costs almost nothing in quality and multi-query attention costs something measurable. What is the mechanism behind the difference?
answer sketch
Sharing a key-value head forces every query head in the group to attend over the same keys and retrieve from the same values, so the group's heads can differ only in how they weight a shared basis. With one shared pair the whole layer loses the ability to have heads looking at genuinely different things, which shows up as instability on tasks that need several distinct retrievals at once. With eight groups there is enough diversity left that the summarisation and QA numbers in the GQA paper sit within a few tenths of multi-head. The uptraining recipe matters too: mean-pooling the group's projections is what makes the converted checkpoint recoverable at five percent of pretraining compute.
-
A sliding window bounds the cache. Why does it collapse without attention sinks, and what does that tell you about the softmax?
answer sketch
The softmax is forced to distribute all of its mass over whatever is in the window, even when nothing there is relevant. A trained model exploits the first few tokens as somewhere to put the mass it does not want, because every position can see them, so evicting them changes the scale of every attention output at once and perplexity explodes. Pinning four of them restores it. The lesson is that softmax has no null option, and that some of what looks like attention is really normalisation bookkeeping.
-
Swapping a preempted sequence to host memory is faster in wall clock than recomputing its prefill, and vLLM recomputes anyway. Defend the choice.
answer sketch
The two costs are paid in different currencies. Swapping spends a shared PCIe link and needs a second memory tier with pinned buffers and its own scheduling; recompute spends tensor-core time that a memory-bound decode step was wasting anyway, and it can be chunked into batches that are already running so it never stalls anyone. It also warms the prefix cache. The engineering argument, that there is one memory tier instead of two, is worth as much as the timing.
-
When is prefix caching worthless, and when is it the entire system?
answer sketch
Worthless when prompts share no exact prefix, because a block is reusable only if every token before it matches, so a personalised token at the front invalidates everything after it. It is the whole system in agent loops and document question answering, where the same long preamble or document precedes every turn and the hit rate reaches the levels SGLang measured in production. That asymmetry is also a prompt-engineering instruction: stable content first, variable content last.
-
Sparse attention divides the bytes read per step by sixty-four at 128k and does not reduce the memory at all. Is that a good trade?
answer sketch
It depends on which limit you are against. If the pool is full and you are turning users away, sparse reading buys you nothing, and you need fewer bytes per token instead. If you have the memory but every step is crawling, it is the largest single lever available, and it leaves quality closer to the dense model than a window does because nothing is permanently forgotten. Also worth noting that the indexer still scores every earlier token, so the cost is linear in the context rather than constant.
-
Predict what speculative decoding does in three settings: batch 1 at 4k, batch 256 at 4k, and batch 8 at 128k. Explain the pattern.
answer sketch
At batch 1 it is close to the full accepted-run multiplier, because the step is deeply memory bound and the extra arithmetic is free. At batch 256 the verify pass, with several times the query positions, has become compute bound and the gain shrinks toward the ratio of memory time to compute time, possibly below one once the drafter's own cost is counted. At batch 8 and 128k the cache is so large that the step is still firmly memory bound, so it pays again. The pattern is that speculation helps exactly while the step is memory bound, and long context keeps it there.
Level 3: can you defend a design decision?
-
You have one node of eight H100s and a product requirement of 128k context for fifty concurrent users. Argue for an approach, and say what measurement would tell you it had failed.
answer sketch
The arithmetic says a dense BF16 cache gives you ten users, so something has to give and the honest first move is to say which. An FP8 cache doubles it; a model with latent attention multiplies it again; a bounded window would get you there outright but changes what the product can do, so it is a product decision, not an infrastructure one. If none of that is enough, the answer is more GPUs or a smaller promised context. Failure shows up as preemption rate and queue depth long before it shows up as latency, so those are what I would put on the dashboard, along with an evaluation on real long-context traffic rather than perplexity.
-
Would you ship a 4-bit KV cache in 2026? Defend either answer.
answer sketch
Not by default. The public evidence that inference teams rely on stops being comfortable below eight bits: FP8 has a production track record with a known, small regression on hard reasoning, and the sub-8-bit results come with layout requirements, per-channel keys and per-token values, and error that concentrates on the attention sinks that every query reads. The case for shipping it is workload-specific: if the traffic is long-context retrieval where the alternative is turning users away, and you can measure quality on that exact traffic, four bits with sink-aware handling is a reasonable experiment. What I would not do is treat a perplexity number as evidence for a reasoning workload.
-
Should the KV cache be a cluster-level resource with its own storage tier, or should it stay per-GPU? What would decide it for you?
answer sketch
The hit rate decides it. A tiered cache turns a prefill into a transfer, which is a large win when many requests share long prefixes, and pure overhead when they do not, so the first thing to measure is the prefix hit rate on real traffic. The second is the fabric: the whole architecture assumes moving a gigabyte is cheaper than recomputing it, which is true over RDMA and false over ordinary networking. And it adds a routing problem, because a request now wants to land where its prefix already lives.
-
Given a free choice of model, would you rather have latent attention, or grouped-query attention with an FP8 cache and sparse reading?
answer sketch
They are not exclusive, and the best answer says so: latent attention is a training-time decision about resident bytes, and sparse reading is an inference-time decision about traffic, so a model can have both and the strongest 2026 systems do. If forced to choose, it depends on which limit binds. Latent attention if capacity is the wall, because it shrinks what has to be resident. Sparse reading if the pool is fine and every step is slow. The risk with latent attention is ecosystem: kernels, quantisation support and serving features arrive later for architectures fewer people run.
-
Prefill and decode disaggregation is the default in large deployments and a mistake in small ones. Where is the line, and what would you measure to find it?
answer sketch
The line is where the coordination and transfer cost stops being small next to the prefill it moves. With a fast fabric a prompt's cache crosses in a few percent of the time its prefill took, and the win is that decode users stop waiting behind prefills. At low load there is nothing to separate: the pools sit idle in turn, and the extra hop is pure latency. I would measure the fraction of decode steps delayed by a prefill in the shared setup, and the transfer time as a fraction of prefill time in the disaggregated one; if the first is small, disaggregation is solving a problem you do not have.
-
Context windows keep growing. What breaks first as they head toward ten million tokens, and which of the current techniques survives that far?
answer sketch
Capacity breaks first and hardest, because resident bytes are linear in the context and no amount of sparse reading changes that. Techniques that only reduce traffic keep the model fast and leave you unable to hold a single sequence. What survives is anything that bounds or compresses what is stored: latent representations, aggressive quantisation, tiered storage where cold blocks live off the GPU, and windows or hierarchical schemes that accept forgetting. The honest position is that ten million tokens of exact recall is a storage architecture problem, not an attention problem.
-
Two users of your API send the same system prompt and share cache blocks. What is the argument that this is a security problem, and how would you bound it?
answer sketch
The shared state is observable through timing: a request that hits a cached prefix has a visibly shorter time to first token, so a user can test whether some other user has recently sent a particular prefix, which leaks content rather than weights. The bounds are all coarse: scope the cache per tenant or per API key, exempt anything user-supplied from sharing and share only prompts the operator installed, or add noise to the latency, which costs the thing you built the cache for. It is a genuine trade between a large efficiency win and a small, real information leak, and the right answer depends on whether tenants are mutually trusted.
-
Is the KV cache fundamental to autoregressive generation, or an artifact of attention? What does the alternative give up?
answer sketch
An artifact of attention specifically. Recurrent and linear-attention families carry a fixed-size state instead, so decoding costs the same per token whatever the history, which removes every capacity and bandwidth problem in this series at a stroke. What they give up is exact recall: a constant-size state cannot hold an arbitrary prefix, so retrieval from far back degrades in a way attention's does not. That is why the interesting 2026 architectures are hybrids that interleave a few full-attention layers with cheap ones, and why the design question has become how few exact-recall layers you can get away with rather than whether to have them.
What should you actually walk in with?
If you remember four numbers you can rebuild most of this from first principles: 320 KiB per token for a Llama-70B-shaped model in BF16, 141 GB for its weights, about 1.3 million cached tokens in a node of eight H100s, and 295 FLOPs per byte for the H100’s ridge, which decode never approaches.
And if you remember one sentence, make it the one that connects them. A decode step is bytes over bandwidth, the cache is most of the bytes, and every technique in this series either shrinks it, shares it, reads less of it, or gets more tokens out of each read.
That framing is what the good answers in this question set have in common. It is also the honest description of what serving work is: not a bag of tricks, but one bottleneck, examined until it becomes obvious what to do about it.