Can you answer these without looking anything up?

Six parts of derivation, and the only way to find out whether it stuck is to be asked. So here are the questions, graded, with the answer I would want to hear rather than a full essay.

Level 1 checks whether you can state the mechanism: shapes, collectives, what is stored, where the bytes go. Level 2 checks whether you can reason about a trade-off you have not seen before, with arithmetic you can do on a whiteboard. Level 3 is the frontier: decisions real teams argued about, where the honest answer is that it depends, and the interesting part is on what.

Read the question, answer it out loud, then open the sketch. Saying it out loud is the whole point: the failure mode under an interview clock is not that you do not know the mechanism, it is that the nuance slips on the way from your head to your mouth.

They will ask Tell me about a time you had to choose a parallelism configuration.
The 30-second version The shape of a good answer here is not a war story, it is a derivation with a decision at the end. Say the model and the cluster. Say the byte inventory: sixteen bytes per parameter for the states, thirty-four times sequence times hidden per layer for the activations. Say what did not fit and by how much. Then say which axis you reached for first and why that one rather than another, with the collective it cost and the link it rode on. Then say what you measured, what surprised you, and what you would do differently. If nothing surprised you, you did not measure enough.

Level 1: can you state the mechanism?

level 1the mechanism
  1. A ring all-reduce runs over d ranks on a message of M bytes. How many bytes does each rank send, and which of bandwidth cost and latency cost grows with d?

    answer sketch

    Each rank sends two times d minus one, over d, times M, which is always less than 2M and does not grow with the ring size. That is the bandwidth cost, and it is flat. The latency cost is not: there are two times d minus one sequential hops, so a large ring pays many link latencies before the last byte lands. That is why libraries switch to tree or hierarchical algorithms for small messages and keep rings for large ones.

  2. Write the per-GPU model-state memory for ZeRO stages 1, 2 and 3 in terms of the parameter count and the data-parallel degree, and name the terms that never shrink.

    answer sketch

    Stage 1 is four psi plus twelve psi over d; stage 2 is two psi plus fourteen psi over d; stage 3 is sixteen psi over d. The four psi in stage 1 is the replicated BF16 weight and BF16 gradient. The two psi in stage 2 is the replicated BF16 weight alone. Those are floors: no amount of data parallelism gets past them, which is the entire reason stage 3 exists.

  3. In a Megatron-style transformer block, which weight matrix is split along its columns and which along its rows, and what forces that order rather than the reverse?

    answer sketch

    The first matrix of a pair is column-parallel and the second is row-parallel. The nonlinearity forces it: a column split gives each GPU whole output features, so anything elementwise applies locally, while a row split gives partial sums that a nonlinearity cannot be applied to without summing first. Doing it the other way round would put a synchronisation in the middle of the block instead of at the end of it.

  4. Define the operators usually called f and g in tensor parallelism, and say what each does in the forward and backward pass.

    answer sketch

    f sits at the entry to the parallel region: identity in the forward pass, all-reduce in the backward, because every GPU produced a gradient with respect to the same broadcast input. g sits at the exit: all-reduce in the forward pass, identity in the backward, because the output was a partial sum and the incoming gradient is already identical everywhere. They are conjugates, and each is a few lines of autograd.

  5. How many collectives does tensor parallelism run per transformer layer per micro-batch, and what is the size of each message?

    answer sketch

    Four: one after the attention block and one after the MLP in the forward pass, and the two conjugates in the backward. Each message is sequence times micro-batch times hidden elements, and a ring moves two times t minus one over t of that per rank. The important part is what it is per: this happens every micro-batch, not every optimizer step, so gradient accumulation does not amortise it at all.

  6. Write the pipeline bubble fraction and say exactly what is in the numerator and the denominator.

    answer sketch

    It is p minus one over m. The numerator is the warm-up plus the drain, each of which is p minus one slots of a forward and a backward, because a stage cannot start until the micro-batch has traversed the stages before it. The denominator is m slots of useful work. It is a ratio to the ideal time, so a bubble of 20% means the step takes 1.2 times the ideal, not that 20% of the step is idle.

  7. Under a 1F1B schedule, how many micro-batches of activations does stage i hold, and how many layers' worth does the first stage hold in total?

    answer sketch

    At most p minus i micro-batches on stage i, so p on the first stage, and that is what caps it independently of the batch. Each of those carries L over p layers of activations, so the first stage holds L layers' worth in total, for every p. Deepening the pipeline divides the weights and does not thin the activations at all.

  8. Name the two all-to-alls in a mixture-of-experts layer and say what each one carries.

    answer sketch

    The dispatch carries each token's hidden vector to every GPU hosting one of its chosen experts. The combine carries the expert outputs back to the GPU the token came from, to be weighted and summed. Both run again in the backward pass, so four per layer per micro-batch, and the volume is tokens times how many places each token has to travel to times the hidden size.

Level 2: can you reason about the trade-offs?

level 2the trade-offs
  1. You double the data-parallel degree at a fixed global batch. What happens to the all-reduce time, the compute per rank, and the exposed fraction of the step?

    answer sketch

    The all-reduce time barely moves, because per-rank volume is flat in the ring size and only the latency term grows. The compute per rank halves, because the tokens per rank halved. So the exposed fraction roughly doubles, and the throughput per GPU falls even though nothing about the model changed. That is exactly the effect Meta reported when Llama 3's data-parallel degree went from 64 to 128 and MFU fell two points.

  2. ZeRO-3 moves 1.5 times the elements of ZeRO-2. When would you deliberately choose stage 2, and what did Meta actually do for Llama 3?

    answer sketch

    When memory is not the binding constraint, because stage 3's extra volume buys nothing you need. Meta took a middle position: they sharded optimizer states and gradients but did not reshard the parameters after the forward pass, so the backward pass reuses the gathered weights and the second all-gather disappears. That is stage 3's forward behaviour with stage 2's traffic, and it is the right answer whenever one unit's gathered weights fit comfortably.

  3. Sequence parallelism replaces four all-reduces per layer with four all-gathers and four reduce-scatters. Why is that not twice the communication?

    answer sketch

    Because a ring all-reduce already is a reduce-scatter followed by an all-gather. Splitting the two halves apart and putting each at a different boundary of the parallel region moves the same bytes in the same number of hops. What it buys is that the norms, the dropouts and the residual stream are now split along the sequence, so the ten-times-sbh term that tensor parallelism could not divide finally divides.

  4. Someone proposes raising the tensor-parallel degree from 8 to 16 on nodes of eight GPUs. Give three separate things that get worse.

    answer sketch

    The group now straddles two nodes, so 320 collectives per micro-batch move from a 450 GB/s link to a 50 GB/s one while the arithmetic they overlap with has halved, taking the ratio from about 38% to well over 100%. The model has eight key-value heads, so sixteen-way splitting duplicates them and adds a reduction to keep the copies in step. And the FFN matmul narrows, which costs efficiency inside the kernel for the reasons the roofline gives.

  5. GPipe and 1F1B have identical wall clocks on the same configuration. So what does 1F1B buy, and how does that turn into a faster run?

    answer sketch

    It caps the micro-batches in flight at the pipeline depth instead of the batch length, so the first stage holds p of them rather than all m. That makes a large m affordable, and a large m is what shrinks the bubble. So 1F1B does not make any given configuration faster; it makes the configuration that is faster possible.

  6. Interleaving divides the bubble by v and multiplies the point-to-point messages by v. When is that a bad trade?

    answer sketch

    When the pipeline dimension is stretched across an oversubscribed tier of the network, so the extra messages contend with data-parallel traffic, or when the sends are already exposed rather than hidden behind compute. It is also bad when the layers do not divide evenly into vp chunks, since an unbalanced chunk becomes a straggler at every stage. On a well-provisioned fabric with sixteen-megabyte messages it is close to free, which is why Llama 3 used it.

  7. Pipeline parallelism divides the weights by p. Why does it not divide the activations, and what would have to change for it to?

    answer sketch

    Because a bubble-minimising schedule needs p micro-batches in flight to keep every stage busy, and each carries L over p layers, so the product is L regardless. To make it divide you would have to keep fewer micro-batches in flight, which reintroduces the bubble, or recompute rather than store, which spends compute. The one exception is that recomputation combines well here, since the stage only has to regenerate its own layers.

  8. Ring attention across a 400 Gb/s network needs roughly six figures of tokens per rank before the transfer hides. Given that, when would you still choose a ring over an all-gather?

    answer sketch

    When the ring stays inside an NVLink domain, where the required block is an order of magnitude smaller, or when the keys and values are large relative to the queries so an all-gather would be expensive. Grouped-query attention pushes the decision the other way, because it makes K and V small, which is exactly the reason Meta chose an all-gather for Llama 3 and accepted the exposed latency in exchange for supporting arbitrary attention masks.

Level 3: can you defend a design decision?

level 3the frontier
  1. DeepSeek-V3 trained a 671 billion parameter model with no tensor parallelism at all. Defend that choice, and say what it cost them.

    answer sketch

    Their expert all-to-all was already using the node's bandwidth, so adding four tensor collectives per layer on the same links would have competed with it directly. They bought the memory back with aggressive recomputation, keeping the exponential moving average on the host, and sharding the states with ZeRO-1 instead. The cost is that every activation inside a block stays full width, which constrains the micro-batch, and that they had to build a bidirectional pipeline schedule and hand-written communication kernels to make the remaining collective disappear. It is a defensible trade only because the model is sparse: a dense model of that size has nowhere else to put the parameters.

  2. 512 H100s, a 70 billion parameter dense model, a global batch you are told not to exceed. Choose every degree and defend it.

    answer sketch

    Tensor 8 with sequence parallelism, because it is the only axis that divides weights and activations together and eight is what a node holds and what the key-value head count allows. Fully sharded data parallelism across the 64 nodes, because its collective fires once per step and prefetches. No pipeline, because the states already fit at that point and a pipeline would cost a bubble and a larger blast radius for nothing. Micro-batch chosen so the activations leave headroom, and gradient accumulation to reach the batch. Then measure, because the model above is a back-of-the-envelope one and the kernels have opinions of their own.

  3. Zero-bubble scheduling gets the last of the bubble by removing the synchronisation in the optimizer step. What is risky about that, and how would you convince yourself the run is still correct?

    answer sketch

    The synchronisation exists so that every stage agrees on things like the global gradient norm before anyone updates, and skipping it means a stage can start the next step on an assumption. The paper's approach is to proceed optimistically and validate afterwards, rolling back if the assumption was wrong. I would want a bit-wise comparison against a synchronous run for a few hundred steps, a check that the skip-and-clip logic still fires identically, and a monitor on how often the rollback path is taken, because a rollback that never fires in testing and fires under a loss spike is the worst version of this.

  4. For expert balance: a capacity factor drops tokens, an auxiliary loss distorts the objective, a routing bias is a non-differentiable control loop. Which do you ship?

    answer sketch

    The bias, with a small capacity factor as a safety net, and I would say why rather than just naming it. An auxiliary loss is a gradient pulling against the language-modelling objective for the whole run, and its weight is one more thing to tune against quality. A bias changes the routing decision without entering the loss at all, which is what DeepSeek-V3 did. The risk is that it is a control loop with a rate constant, so it can oscillate or lag a distribution shift, and it needs monitoring on per-expert token counts rather than trust.

  5. Your team wants to go from 512 GPUs to 4096 at a fixed global batch. What breaks first?

    answer sketch

    The tokens per rank, and everything downstream of it. At fixed batch, eight times the data parallelism means an eighth of the tokens each, so the fixed data-parallel collective has an eighth of the compute to hide behind, and if a pipeline is involved the micro-batch count falls by the same factor and the bubble grows. The honest options are to raise the global batch and re-tune the learning rate, or to spend the extra GPUs on model-parallel axes rather than data-parallel ones. It is also where the failure arithmetic starts to bite, because eight times the GPUs is roughly an eighth of the mean time between restarts.

  6. Argue for and against setting the tensor-parallel degree to the number of key-value heads rather than the number of GPUs in a node.

    answer sketch

    For: beyond the key-value head count the heads have to be duplicated, which wastes memory and adds a reduction, so the head count is a genuine architectural ceiling and it is often the smaller of the two. Against: the degree also has to divide the node cleanly, or the group straddles a node boundary and the collectives land on the slow link, which is a much larger effect than a duplicated key-value head. When the two disagree, the topology wins and the architecture should be changed instead, which is one reason models are designed with head counts that are powers of two.

  7. Blackwell systems put 72 GPUs in a single NVLink domain. What does that change about the rules in this series, and what does it not change?

    answer sketch

    It moves the boundary, not the rule. The rule is that the highest-frequency collective goes on the fastest link, and a 72-GPU domain means the tensor-parallel degree is no longer capped at eight by the topology. What still caps it is the architecture, since the key-value head count and the GEMM width do not change, and the ratio argument still applies inside the domain, since communication climbs towards a ceiling while compute falls as one over the degree. What it most changes is expert parallelism, whose all-to-all is the collective that suffers most from crossing a network, and which now has a much larger domain to stay inside.

  8. When would you accept a lower model FLOPs utilisation on purpose?

    answer sketch

    Whenever the thing you are optimising is not GPU-hours. Long-context training is the clearest case: context parallelism cost Llama 3 several points of MFU and bought a capability that no amount of throughput would have. Reliability is another, since a shallower pipeline has a smaller blast radius and restarts less often, and a run that is 5% slower and restarts half as much finishes sooner. And recomputation is the everyday case: it spends compute the MFU refuses to count, in exchange for a micro-batch large enough that everything else runs better.

That closes the series. Two questions started it: what does one GPU have to hold, and how long would it take.

Six parts later the answer is a list of five axes, a table of what each divides, a collective attached to each, and one rule for mapping them onto a machine: sort the collectives by how hard they push, sort the links by how much they carry, and pair them off. Everything else is arithmetic, and arithmetic is the part that survives an interview clock.