What actually goes wrong at scale?
Six parts of theory, and every one of them survives contact with a real cluster. What does not survive is the assumption that the cluster behaves like one big GPU. It is sixteen thousand of them, each with a tenth of a percent chance of failing this week, all waiting for the slowest one at every synchronisation, fed by a file system and a data loader that were never in the FLOP count.
This part is the list I wish someone had handed me before my first large run: how to measure whether the machine is being used, where the unmeasured time goes, how often the hardware will fail and what that implies for checkpointing, what a loss spike is and how to survive one, and what the host has to do while the GPUs work. It ends with a checklist, because that is how it gets used.
The numbers are from Meta’s Llama 3 report where a public number exists, and from the earlier parts where the arithmetic is ours.
The 30-second version
Measure the right thing: model FLOPs utilisation, six FLOPs per parameter per token achieved over peak, which a healthy frontier run puts at about 40%, and be explicit about whether recompute is counted. Account for the other 60%: kernels below the roof, exposed communication, pipeline bubbles, recompute, stragglers at every synchronisation, and host stalls. Budget for failures: per-GPU reliability of years becomes a cluster failure every few hours at ten thousand GPUs, so checkpoint at the Young-Daly interval, asynchronously through host memory, to a separate failure domain. Watch the numerics: gradient norms and activation statistics every step, QK-norm or QK-clip against logit growth, fine-grained scales and FP32 accumulation in low precision, and a rewind-and-skip procedure for spikes. Keep the host ahead: tokenised, sharded, prefetched data in pinned memory, and launches reduced with CUDA graphs. And fix the allocator before adding GPUs, because fragmentation at the peak looks exactly like being out of memory.How do you measure whether the machine is being used?
Tokens per second is the number people quote and the wrong one to optimise, because it says nothing about how much of the machine did the work. The right number divides the FLOPs the model needs by the FLOPs the GPUs could have done:
\[\text{MFU} = \frac{6 P \cdot \text{tokens per second}}{N_{\text{GPU}} \cdot \text{peak FLOP/s}}\]The six is two FLOPs per parameter per token in the forward pass and four in the backward. Attention adds $12 L h s$ per token on top, quadratic in the sequence, which most reports leave out of the numerator; when comparing numbers, ask. Hardware FLOPs utilisation counts recomputed activations as useful work and reads about ten points higher than MFU for the same run; ask about that too.
Llama 3 405B reported 38 to 43% BF16 MFU on 16,384 H100s. That is what a well-tuned frontier run looks like, and the figure’s first step starts there. Above 50% is not a real number for a transformer at that scale, and a report claiming it is usually counting something else.
Where does the other sixty percent go?
Take the 40% and account for the rest, because the rest is where the engineering is.
The matmuls themselves rarely exceed 75% of peak: tiles are finite, epilogues are not free, and the L2 does not always help. Attention and every non-matmul kernel run below the roof for the reasons in Part 2. That is the first fifth of the missing time, and it is the kernel engineers’ fifth.
Communication that did not hide under compute is the next slice: the all-gathers and reductions that stuck out past the end of the backward pass. Pipeline bubbles, the warm-up and drain of each pipeline, are another. Activation recomputation spends GPU time on work the MFU refuses to count. And at every synchronisation, the whole job waits for its slowest GPU, so a single throttling part or a slow link taxes sixteen thousand.
The last slice is the host: data loading, kernel launches, the pause for checkpoints, and the restart after each failure. The figure’s second step draws a plausible split for a healthy run. The shares are an illustration, not a measurement; the shape is what to learn. The profiler tells you your own split in an afternoon, and every slice maps to one of Part 1’s clocks.
How often will the hardware fail?
Per GPU, almost never. Per cluster, constantly.
Meta’s Llama 3 report is the public benchmark. In a 54-day window on 16,384 H100s there were 466 interruptions, 419 of them unexpected, one every three hours. Faulty GPUs caused 148 of them and their HBM3 memory 72; GPU SRAM 19, the GPU’s system processor 17, network switches and cables 35. Seventy-eight percent of the unexpected interruptions were hardware. Three needed a human; automation handled the rest.
Divide it out and the mean time between failures is roughly 50,000 hours per GPU, nearly six years. At 16,384 GPUs that is three hours for the job, and every GPU you add shortens it. This is the arithmetic that turns reliability from an operations topic into a design constraint.
The checkpoint interval follows from it. Checkpoint too often and you pay the pause; too rarely and you pay to redo the work since the last one. Young and Daly’s formula balances the two:
\[T_{\text{opt}} = \sqrt{2 \cdot C \cdot \text{MTBF}_{\text{cluster}}}\]where $C$ is the time the job pauses to checkpoint. With a three-hour cluster MTBF and a thirty-second pause, the optimal interval is about fourteen minutes, and the run loses roughly 6% of its time to writing and redoing. Halve the pause and both numbers improve, which is why the pause is the thing to engineer.
What is a loss spike, and how do you survive one?
The other failure is silent. The loss jumps, or the gradient norm does, or a NaN appears in an FP8 tile, and sixteen thousand GPUs keep running while the model unlearns.
The causes are almost all numerical, and the earlier parts pointed at each of them. Attention logits grow until the softmax saturates, which QK-norm during training or QK-clip after each step prevents; Kimi K2 reported zero spikes over 15.5 trillion tokens with the latter. Low precision lets one outlier take a whole tensor with it unless the scales are fine-grained. Gradient accumulation across micro-batches loses small contributions unless the accumulator is FP32.
The defences are boring and non-negotiable. Log the gradient norm and per-layer activation statistics every step. Skip or clip a step whose gradient norm is an outlier. Keep a recent checkpoint you can rewind to, and a data pipeline you can skip forward in, so a spike costs the steps since the last checkpoint rather than the run; OPT, PaLM and Llama 3 all documented doing exactly this. And when the precision recipe changes, run a BF16 control long enough to compare loss curves before trusting the faster one.
One more, from the same family: silent data corruption. A GPU can compute the wrong answer without raising an error, and at scale one eventually does. The check is redundancy, the same batch on two GPUs should agree, run occasionally as a health test.
What does the host have to do?
The GPU is not the only machine in the room, and two host-side jobs decide whether the GPUs wait.
The data loader has to deliver tokens faster than the GPUs consume them. That means tokenisation done offline, files sharded so every rank reads its own, prefetching with worker processes, and the next batch already in pinned host memory when the step begins. A stalled loader is a stalled cluster, and in a summary it looks exactly like slow GPUs. Measure the loader alone before blaming anything else.
Checkpoints have to be written without stopping the job. The shape every framework converged on: pause, copy each GPU’s shard of the states to pinned host memory, resume, and let a host thread stream the copy to storage while training continues. PyTorch’s distributed checkpoint does this with asynchronous saves. The pause $C$ in the Young-Daly formula becomes the copy, seconds, rather than the write, minutes. Meta’s file system sustained 2 TB/s for Llama 3, which makes even a 6.5 TB checkpoint of the 405B model’s full states a few seconds of aggregate write. Write to a different failure domain than the one you run in, and verify a checkpoint before deleting the one before it.
The launch queue from Part 1 is a host problem too, and CUDA graphs and compilation solve it on the host’s behalf. Below a few hundred tokens per step they are the whole optimisation.
What do you check before trusting a number?
In order, each one a part of this series folded into a sentence.
- Fit before fast. Write the byte inventory: sixteen bytes per parameter for states, the activation formula, the KV cache, the overheads. Then choose what to divide by what.
- Measure MFU, not tokens per second. Say whether it counts recompute. Compare against 40% for a healthy frontier run.
- Put the kernel on the roofline. FLOPs, bytes, intensity, ridge. A kernel at 5% of peak may be at 100% of bandwidth.
- Find the clock. Gaps in the timeline mean launches. DRAM-bound kernels mean bytes. Exposed NCCL means communication.
- Overlap or shard, never sum. Communication hidden behind the backward pass costs nothing. Exposed, it is the step.
- Change one precision at a time, with a control. Fine-grained scales, FP32 accumulation, a BF16 run to compare against.
- Budget the failures. MTBF per GPU over the cluster size is the job’s lifetime. Checkpoint at the Young-Daly interval, asynchronously, elsewhere.
- Watch the norms. Gradient norm and activation statistics every step. A spike caught in ten steps costs ten steps.
- Fix the allocator before adding GPUs. Reserved far above allocated is fragmentation. Expandable segments first.
- Small batch means launch-bound. Below a few hundred tokens per step, CUDA graphs and fusion are the whole optimisation.
The 30-second version
First confirm the denominator and the numerator: dense BF16 peak for the GPU, six FLOPs per parameter per token, and whether recompute is on, because full recompute alone costs a third and would explain most of the gap. Then profile one step on a few ranks. Read the GPU timeline for the four signatures: gaps between kernels are launches or a starved data loader; NCCL kernels not overlapped with compute are exposed communication, so check the bucket sizes and the FSDP prefetch; a long idle at the pipeline ends is bubbles, so check the micro-batch count against the pipeline depth; and if one rank is consistently later than the others at each synchronisation, it is a straggler, a throttled GPU or a bad link. Then look at the kernels themselves: is attention on FlashAttention-3 or 4, are the norms fused, are the matmul shapes multiples of the tile size. Most 28% runs turn out to be exposed communication plus a recompute setting nobody revisited after the sequence length changed.Rapid fire: can you do these from memory?
- Write the MFU formula and say what the six counts and what it leaves out.
- Name the six places the non-MFU time goes and the profiler signature of each.
- From 419 failures in 54 days on 16,384 GPUs, derive the per-GPU MTBF and the cluster MTBF.
- Derive the Young-Daly checkpoint interval and evaluate it for a three-hour cluster MTBF and a 30-second pause.
- Name three numerical causes of loss spikes and the defence against each.
- What is asynchronous checkpointing, and which term in the checkpoint formula does it shrink?
- How would you tell a slow data loader from a slow GPU in a profile?
- Why does fragmentation look like an out-of-memory error, and what is the one-line fix?
That closes the series. Two questions started it: where does the time go, and where do the bytes go. Seven parts later the answer is a short list of clocks, a chart with a corner, a pyramid, an inventory, one kernel that respects all of them, a catalogue of what moves which line, and the operational arithmetic that keeps sixteen thousand GPUs pointed at the same loss. That is what I want to carry into the room.