Why can one GPU not train this model?

The GPU series ended with an inventory: sixteen bytes per parameter, an activation formula, a KV cache, and four clocks that decide which of them is the bottleneck. That inventory is for one GPU. The next question an interviewer asks is the one that follows from it, and it is almost always phrased as a task rather than a question: “you have 512 H100s and a 70 billion parameter model. Lay it out.”

There are five ways to cut a model across GPUs, and they are not interchangeable. Each one divides a different line of that inventory, each one buys the division with a different collective, and each one runs out of road for a different reason. Getting the list right is easy. Knowing which line each one divides, and what it costs, is the part that separates a candidate who has read about this from one who has run it.

This series derives the five, one at a time, on one model and one cluster. Everything is arithmetic you can do on a whiteboard, and every number in every figure is computed from the same constants.

They will ask Why does a large model need more than one GPU, and what are the ways to split it across them?
The 30-second version Two budgets force it, independently. Memory: Adam in mixed precision costs sixteen bytes per parameter, so a 70B model needs 1.13 TB of state before a single activation, and one 8192-token sequence adds another 182 GB of saved activations. That is sixteen H100s' worth on a machine that has 80 GB. Time: six FLOPs per parameter per token over fifteen trillion tokens is 6.3 times ten to the twenty-fourth, which one H100 at a realistic 40% of peak would finish in five centuries. Five axes split it. Data parallelism splits the batch and divides the optimizer states, gradients and weights through ZeRO, paying with an all-reduce or a reduce-scatter and all-gather per step. Tensor parallelism splits the matmuls inside each block and divides the activations too, paying with four collectives per layer, which is why it stays inside a node. Pipeline parallelism splits the layers, paying in idle time rather than bytes. Context parallelism splits the sequence, for long-context training. Expert parallelism splits a mixture of experts, paying with two all-to-alls per layer. Real runs use three or four of them at once, ordered so that the busiest collective gets the fastest link.

What does one GPU actually have to hold?

Take Llama 3 70B, whose shape Meta published: 70.5 billion parameters, 80 layers, hidden size 8192, an FFN width of 28,672, 64 query heads and 8 key-value heads of dimension 128.

Training one parameter with Adam in mixed precision costs sixteen bytes: a BF16 weight, a BF16 gradient, an FP32 master copy, and two FP32 optimizer moments. That is 1.13 TB of model state, and none of it is optional while the optimizer is Adam.

Then the activations. Megatron’s accounting for one transformer layer, with FlashAttention removing the quadratic attention term, is $34\,s\,b\,h$ bytes. At a sequence of 8192 tokens and $h = 8192$ that is 2.28 GB for one layer, and 182.5 GB across all eighty.

Add them and one GPU would need 1.31 TB. An H100 has 80 GB. The first figure draws that as sixteen and a bit GPUs’ worth of boxes, and then empties them one axis at a time.

How long would one GPU take anyway?

The second budget is time, and it does not care about the first. A dense transformer costs about six FLOPs per parameter per token: two in the forward pass, four in the backward.

\[\text{FLOPs} = 6 P D = 6 \times 70.5\times 10^{9} \times 15 \times 10^{12} \approx 6.3 \times 10^{24}\]

An H100 does 989 TFLOP/s of dense BF16 at peak, and a well-tuned training run achieves about 40% of that. So one GPU delivers 396 TFLOP/s of useful arithmetic and finishes the run in roughly five hundred years.

Five hundred and twelve of them take 363 days. Sixteen thousand take eleven. That is the shape of the second budget: it is satisfied only by throwing hardware at it, and the hardware is only useful if the model is split in a way that keeps it busy.

Notice that the two budgets give different answers. Memory says you need at least seventeen GPUs. Time says you want thousands. The interesting engineering lives in the gap.

What are the five things you can split?

Five, and it helps to name them by what gets cut rather than by their acronyms.

Split the batch. Every GPU holds the model and processes different tokens. Gradients are averaged at the end of the step. ZeRO and FSDP are refinements that stop replicating the states.

Split the tensors. Each matrix multiply inside a block is cut across GPUs, so every GPU holds a slice of every layer and works on the same tokens.

Split the layers. Each GPU holds a contiguous group of layers and passes activations to the next group, like a factory line.

Split the sequence. Each GPU holds part of the token sequence, which only matters because attention couples the whole sequence together.

Split the experts. In a mixture-of-experts model, the experts live on different GPUs and tokens travel to them.

Data parallelism is the only one of the five where the model is never cut. That is why it is the default, and why it is the first to run out.

Why does each axis divide a different thing?

This is the table I would draw on the whiteboard before answering anything else, and the fourth step of the figure builds it.

Four things occupy memory during training: the weights, the gradients, the optimizer states, and the activations saved for the backward pass. Every axis divides some of them and leaves the rest alone.

Plain data parallelism divides nothing. ZeRO stage 1 divides the optimizer states, stage 2 the gradients too, stage 3 the weights as well. None of the three touches the activations, because each rank still has to keep the activations of its own tokens.

Tensor parallelism divides all four, which is what makes it valuable and expensive at the same time. Pipeline parallelism divides the three model-state lines by the number of stages, but keeps several micro-batches in flight, so the activations barely move. Context parallelism divides the activations and nothing else. Expert parallelism divides the expert weights and nothing else.

Key idea Every parallelism scheme is a choice of which line of the memory inventory to divide, and the collective it costs is exactly the information you decided not to keep a copy of.

What does each axis cost on the wire?

Numbers, for the layout this series arrives at: tensor parallelism of 8 inside each node, fully sharded data parallelism across the 64 nodes, one micro-batch of 8192 tokens.

One GPU owns 8.81 billion parameters and does 0.44 seconds of arithmetic per micro-batch at peak. Against that:

The gradient all-reduce of plain data parallelism moves 34.7 GB per rank per optimizer step. Fully sharded data parallelism moves 52.9 GB, because stage 3 adds an all-gather in the backward pass. Both ride the 400 Gb/s network at 50 GB/s, so they take of order a second, and both can hide behind the backward pass if the framework overlaps them properly.

Tensor parallelism moves 75.2 GB per micro-batch, not per step, because it fires four times per layer. It gets away with that only because NVLink inside a node is 450 GB/s per direction, nine times the network. On the network the same traffic takes 1.5 seconds against 0.44 seconds of arithmetic.

Pipeline parallelism moves 33.6 MB per stage boundary per micro-batch, forward and backward. It is the cheapest axis on the wire by three orders of magnitude, and it is the only one whose cost is idle time rather than bytes.

What does a good layout look like on this cluster?

The last step of the figure sets it out, and every line in it is derived rather than remembered.

Tensor parallelism of 8, filling one node, divides both the weights and the activations. Fully sharded data parallelism across the 64 nodes divides the states again, so no byte of the model is stored twice anywhere in the cluster. Micro-batch of one 8192-token sequence per tensor-parallel group. Gradient accumulation sets the global batch without changing a single collective.

That gives 2.2 GB of model state per GPU, 22.8 GB of activations, about 4 GB of context and buffers and workspace, and 29 GB of an 80 GB card in use. The arithmetic is 438 ms per micro-batch at peak, and the whole run is 363 days at 40% MFU.

There is a lot of headroom in that 29 GB, and the rest of the series is about what to spend it on: a larger micro-batch, a deeper pipeline, a longer sequence, or nothing at all.

What is this series, and what is it not?

It is the derivation, not the walkthrough. Every part asks the question an interviewer opens with, works out the mechanism from the shapes, and puts a number on it.

For the practical side of the same subject, running jobs on a real cluster, I wrote Distributed Training from Scratch earlier this year: the batch-size hierarchy, DDP against FSDP with real config, and what actually breaks at 256 GPUs. That series is the one to read for launch scripts and the shape of a bad night. This one is the one to read before someone asks you to derive a bubble formula on a whiteboard.

For the hardware underneath, the GPU series has the four clocks, the roofline, the memory hierarchy and the byte-by-byte inventory this part starts from. I will point at it rather than repeat it.

Six parts of derivation, and then a graded question set to check whether it stuck.

They will ask You have 512 GPUs and a 70B model. Where do you start?
The 30-second version With the byte inventory, not with the parallelism. Sixteen bytes per parameter is 1.13 TB of state, and 34 times sequence times hidden per layer is 182 GB of activations for one 8192-token sequence. That is sixteen H100s of memory for one GPU's worth of work, so the model has to be cut before anything else is decided. Then pick the axes in the order the hardware allows: tensor parallelism up to the eight GPUs in a node, because it is the only axis that divides both weights and activations and it needs NVLink; fully sharded data parallelism across the nodes, because its collective runs once per step and hides behind the backward pass. Check the arithmetic: 2.2 GB of states plus 22.8 GB of activations plus overheads is under 30 GB, so it fits with room to raise the micro-batch. Only add a pipeline if the states still do not fit, because a pipeline costs a bubble that data parallelism does not. And say the number: about 363 days on 512 H100s at 40% MFU for a fifteen trillion token run, which is the real reason frontier labs use sixteen thousand.

Rapid fire: can you do these from memory?

  1. Derive the model-state bytes for a 70B model with Adam in mixed precision, and the activation bytes for one 8192-token sequence.
  2. Write the FLOP count for a full training run and evaluate it for 70B parameters and 15T tokens.
  3. Name the five axes and say which line of the memory inventory each one divides.
  4. Which axes leave the activations untouched, and why?
  5. Why does tensor parallelism have to stay inside a node while data parallelism does not?
  6. Estimate the gradient all-reduce volume per rank for 8.8 billion parameters in BF16.
  7. What is the cheapest axis on the wire, and what does it cost instead?

Part 2 takes the axis everyone starts with, splitting the batch, and follows it until it stops working.