Where does the time go in a training step?

The Transformer series ended with a block I could write from memory. The next thing an interviewer does is put a number on it. “Llama 3 70B, eight GPUs, one training step. How long does it take, and what is it waiting on?”

That question has a precise answer, and the answer is not a number. It is a taxonomy. A GPU is four machines sharing a chip: an arithmetic engine, a memory system, a network port, and a command queue fed by a CPU. Each has its own speed, and any piece of work finishes when the slowest of the four is done with it.

Knowing which of the four is the clock, for the operation in front of you, is the whole discipline. Every optimisation in the rest of this series is a way of moving one of them.

They will ask What are the bottlenecks when training or serving a large model, and how do you tell which one you are hitting?
The 30-second version Four clocks run on every operation. Compute time is FLOPs divided by peak FLOP/s. Memory time is bytes moved divided by HBM bandwidth. Communication time is bytes sent divided by link bandwidth. Overhead is kernel launches times a few microseconds each, plus host-side work. The operation takes roughly the maximum when they overlap and the sum when they do not. Large matmuls are compute-bound. Decoding one token at a time is memory-bound, because every weight is read to do two FLOPs per weight. Data-parallel gradient reduction is communication-bound unless it is hidden behind the backward pass. Tiny batches are overhead-bound. Memory capacity is a fifth constraint that does not slow you down but stops you: about 16 bytes per parameter to train with Adam in mixed precision. You tell them apart with a profiler: gaps between kernels mean overhead, low SM throughput with high DRAM throughput means memory, exposed NCCL kernels mean communication, and achieved FLOP/s near peak means you are finally compute-bound.

What are the four clocks?

Take any operation the GPU runs and write down four durations.

\[t_{\text{compute}} = \frac{\text{FLOPs}}{\text{peak FLOP/s}}, \qquad t_{\text{memory}} = \frac{\text{bytes moved}}{\text{HBM bandwidth}}\] \[t_{\text{comm}} = \frac{\text{bytes sent}}{\text{link bandwidth}}, \qquad t_{\text{overhead}} = \text{launches} \times \text{a few } \mu\text{s}\]

The operation takes about as long as the largest of the four when the hardware can overlap them, and about their sum when it cannot. That single sentence is the model behind every profile you will ever read.

The figure runs it on one layer of Llama 3 70B, a GPU’s share of one training step, forward and backward. The numbers are real: 856 million parameters in the layer, an H100 at 989 TFLOP/s of dense BF16 and 3.35 TB/s of HBM, a 400 gigabit network port. Step through it. A different clock is on top in each regime, and the last step puts all of them on one chart.

There is a fifth constraint that is not a clock. Memory capacity does not make you slower; it makes the job impossible until you shard it. It gets its own part later. The four clocks are about time.

When is the tensor core the clock?

Training, at a healthy batch. The layer’s widest matmul multiplies a $T \times 8192$ activation by the $8192 \times 28672$ FFN weight. That is $2 \cdot T \cdot 8192 \cdot 28672$ FLOPs against roughly $2 \cdot 8192 \cdot 28672$ bytes of weight, because the weight matrix dominates the traffic once $T$ is in the thousands.

At $T = 8192$ tokens the matmul does 3.9 TFLOP and moves about a gigabyte. On an H100 that is 3.9 ms of arithmetic and 0.3 ms of memory traffic. The tensor cores are the clock by a factor of twelve, and HBM idles most of the time.

Divide the FLOPs by the bytes and you get about 3,600 FLOP per byte. Divide the H100’s peak FLOP/s by its bandwidth and you get about 295. An operation is compute-bound when the first number exceeds the second, and this one does by an order of magnitude. Part 2 turns that ratio into a picture.

This is the regime GPUs are designed for. It is also the regime that gets harder to reach every generation, because compute grows faster than bandwidth, and much faster than the other two clocks.

When is HBM the clock?

Serving. To generate one token, the model runs a forward pass for a single row. Every layer reads its 1.7 GB of weights from HBM to perform $2 \times 856$ million FLOPs, two per weight. The arithmetic takes 1.7 microseconds on an H100. The read takes 510 microseconds.

Multiply by 80 layers and one token costs about 42 ms of pure weight traffic if the model fit on one H100, which at 141 GB in BF16 it does not. Split across two GPUs with tensor parallelism, each reads 70 GB per token: 21 ms plus a synchronisation. A B200 with 192 GB and 8 TB/s holds the whole model and reads it in 17.6 ms, so a lone sequence gets about 57 tokens a second.

Batching is the fix, because the weights are read once per step regardless of how many sequences are decoding. Compute grows with the batch, the weight bytes do not, and the two clocks cross at a batch near 295 on an H100. Below that, HBM is the clock. Above it, the tensor cores are.

There is a second stream of bytes that the weight arithmetic hides. Each sequence carries a KV cache, 320 KiB per token for this model, and every decode step reads all of it. At 8k context that is 2.7 GB per sequence per step. Thirty-two sequences read 86 GB of cache per step against 141 GB of weights. At long context the cache, not the weights, is what HBM spends its time on, and that is why Part 6 is full of ways to shrink it.

Key idea Training does thousands of FLOPs per byte and is compute-bound. Decoding does two FLOPs per weight byte and is memory-bound. The same model, the same GPU, a different clock, and the difference is entirely how many tokens share each read of the weights.

When is the network the clock?

Data parallelism. Each GPU computes gradients for the full model on its own tokens, and at the end of the step the gradients are averaged. For 70 billion parameters in BF16 that is 141 GB of gradient per GPU. A ring all-reduce makes each GPU send and receive about twice that, 282 GB.

Inside one node, across NVLink at 450 GB/s per direction on an H100, the reduction takes about 0.6 s. Across nodes, on a 400 gigabit port that moves 50 GB/s, it takes 5.6 s. The 800 gigabit ports on Blackwell systems halve that to 2.8 s.

Set that against the compute. At 8k tokens per GPU the step’s arithmetic is $6 \times 70.5\text{B} \times 8192$ FLOPs, which is 3.5 s on an H100 at peak. The un-overlapped reduction is longer than the compute. A naive data-parallel job over a 400 gigabit network is communication-bound.

Nobody runs it naively. Every framework starts reducing the gradients of layer $L$ while the backward pass is still computing layer $L - 1$, so most of the transfer hides behind arithmetic, and only the part that sticks out past the end of the backward pass costs anything. Larger local batches help for the same reason: they make the compute bar longer than the communication bar. Sharded approaches like FSDP and ZeRO change what is communicated, all-gathering weights and reduce-scattering gradients layer by layer, but they live or die by the same overlap.

When is the launch queue the clock?

Tiny batches. A layer is about ninety kernels, forward and backward: projections, a norm, a softmax, residual adds, activations, and their gradients. The CPU tells the GPU about each one, and the telling costs a few microseconds per kernel before any work starts.

At 8k tokens the layer’s arithmetic is 42 ms and its launches are half a millisecond. Nobody notices. At 128 tokens, a fine-tuning micro-batch, the arithmetic is 0.7 ms, the launches are still half a millisecond, and the weight reads are 1.5 ms. The launches are no longer noise, but for a 70B layer the weights are so large that HBM stays on top.

Launch-bound proper is a small-kernel disease. Take Llama 3.2 1B decoding one token. A layer’s weights are 122 MB, read from HBM in 36 microseconds. The forward pass of that layer is about thirty kernels, which is 180 microseconds of launches. The host is the clock by five to one, and across sixteen layers the model produces about 300 tokens a second where the memory system could deliver over a thousand.

The fix is fewer launches. Fusing neighbouring operations into one kernel is what torch.compile does. Recording the whole step’s launch sequence once and replaying it from the device is what CUDA graphs do. Both are invisible at large batch and decisive at small, and every serving engine ships the second one for exactly this reason.

What about memory capacity?

Capacity is the constraint that does not show up on a profile because the job never starts. Training a parameter with Adam in mixed precision costs 16 bytes: two for the BF16 weight, two for its gradient, and twelve for the FP32 master copy and the two optimizer moments. Seventy billion parameters need 1.13 TB before a single activation is stored. An H100 has 80 GB.

Activations are the other half. With FlashAttention, one layer stores about $34 \cdot s \cdot b \cdot h$ bytes for a sequence of $s$ tokens, Megatron’s accounting for a GPT-style block, which at 8k tokens and $h = 8192$ is 2.3 GB per layer and 182 GB across 80 layers for a single sequence. Without FlashAttention there is an extra $5 \cdot a \cdot s^2 \cdot b$ term for the attention matrices that adds 21 GB per layer.

So the 70B model is sharded across at least sixteen GPUs before anyone asks how fast it runs, and the sharding decides which of the four clocks you meet next. Part 4 does that accounting properly, byte by byte.

How do you tell which one you are hitting?

The profiler tells you, if you know what to look for. I use torch.profiler for the timeline and Nsight Compute for individual kernels, and read them against the four clocks.

Gaps between kernels on the GPU timeline, with the CPU busy, are launch overhead. A kernel whose DRAM throughput is near peak while its tensor core utilisation is low is memory-bound. NCCL kernels that are not overlapped with compute kernels are exposed communication. And when the achieved FLOP/s of the step approaches peak, you are compute-bound and done, because that is the only bottleneck money cannot fix from software.

That last number has a name. Model FLOPs utilisation, the fraction of peak the model’s own arithmetic achieves, is how training runs are judged. Llama 3 405B reported 38 to 43% on 16,384 H100s. Part 7 explains why the rest was lost, and how much of it was recoverable.

They will ask A model is decoding at 20 tokens a second on one GPU. What would you try first?
The 30-second version Check the arithmetic. At batch one, decode reads every weight to do two FLOPs per weight, so the token rate is about bandwidth divided by model bytes, and 20 tokens a second says the GPU is doing exactly that. Nothing about the compute matters yet. The levers are all on the byte side: batch more sequences so each weight read is shared, quantise the weights so there are fewer bytes to read, and shrink or quantise the KV cache, which at long context is read every step too. Speculative decoding attacks it from the other side by turning one weight read into several tokens. Only once the batch is in the hundreds does compute enter the picture.

Rapid fire: can you do these from memory?

  1. Write the four durations and say when an operation takes their maximum and when their sum.
  2. Why is a large matmul compute-bound and single-token decode memory-bound, for the same weights?
  3. What is the ratio you compare against a machine's peak FLOP/s divided by its bandwidth, and roughly what is that number for an H100?
  4. How many bytes does a ring all-reduce move per GPU for 141 GB of gradients, and how long is that over 400 gigabits?
  5. Why does communication stop being the bottleneck when it is overlapped with the backward pass?
  6. Estimate the launch overhead of a 70B forward and backward pass. When does it matter?
  7. Why is memory capacity not one of the four clocks?
  8. Name the profiler signature of each of the four bottlenecks.

Part 2 takes the ratio from this part, FLOPs per byte, and puts it on a chart with two lines: the roofline.