Where can a byte live on a GPU?

The roofline counts bytes to and from HBM as if HBM were the only memory. It is the only memory that matters for the roofline because it is the slowest memory on the chip, and everything faster sits above it in a pyramid that a kernel author manages by hand.

That pyramid is where fused kernels come from. FlashAttention does the same arithmetic as naive attention; it just keeps the score matrix on a level of the pyramid that the roofline does not count. To see why that is possible, and why it is hard, you have to know what the levels are, how much they hold, and how fast they move.

Interviewers ask about this in two ways. The soft version is “what are the components of GPU memory?” The hard version is “why does a matmul use 128 by 128 tiles?” They are the same question.

They will ask What are the components of a GPU's memory, and which parts of a model live in which?
The 30-second version From fast to slow: the register file in each streaming multiprocessor, 256 KB per SM, which holds the operands and accumulators of the arithmetic in flight; shared memory, up to 227 KB per SM, a scratchpad the kernel fills explicitly with the tiles it is working on; the L2 cache, 50 MB on an H100 and 126 MB on a B200, shared by all SMs and holding whatever they touched recently; and HBM, 80 to 192 GB at 3.35 to 8 TB/s, which holds everything that has to survive between kernels. Below HBM come other GPUs over NVLink at 450 to 900 GB/s per direction, the host over PCIe at 64 GB/s, and the network at 50 to 100 GB/s. Weights, optimizer states, saved activations and the KV cache live in HBM. A matmul's current tiles live in shared memory and its accumulator in registers, or in Blackwell's 256 KB tensor memory per SM. Attention keeps a block of Q and the streaming blocks of K and V in shared memory and the running softmax statistics in registers. Nothing above HBM outlives the kernel that put it there.

What are the levels?

Every place a byte can live, per GPU, with the H100’s numbers:

level size bandwidth roughly how far
registers 256 KB per SM, 33 MB total about ten times shared memory 1 cycle
shared memory / L1 up to 227 KB per SM, 30 MB total 128 B per clock per SM, 33 TB/s aggregate tens of cycles
L2 cache 50 MB 5.5 to 12 TB/s measured hundreds of cycles
HBM3 80 GB 3.35 TB/s 600 cycles
NVLink to the node’s other GPUs 7 more GPUs 450 GB/s per direction a microsecond
host memory over PCIe 5 1 to 2 TB 64 GB/s per direction microseconds
network the cluster 50 GB/s on a 400 gigabit port several microseconds

The B200 keeps the same shape with bigger numbers below the SM: 148 SMs, a 126 MB L2, 192 GB of HBM3e at 8 TB/s, 1.8 TB/s of NVLink. Inside the SM it adds one level, a 256 KB tensor memory wired into the tensor cores to hold accumulators, and leaves shared memory unchanged at 227 KB and 128 bytes per clock. That last fact is the one to remember.

How steep is the ladder?

Roughly a decade per rung. Shared memory streams about ten times what HBM does in aggregate, 128 bytes per clock in each of 132 SMs at 1.98 GHz, which is 33 TB/s against 3.35. The register file feeds the arithmetic units at roughly ten times that again.

Below HBM the drop is steeper. NVLink carries a seventh of HBM per direction, PCIe a fiftieth, the network a seventieth. Four orders of magnitude separate the register file from the network, in bandwidth and in latency both.

That is why HBM bandwidth is the right number for the roofline. It is the last fast memory before the cliff and the first slow one after the tensor cores. A kernel that goes to HBM for every operand runs on the slope; one that goes there once per tile runs on the roof.

What does one tile’s journey look like?

Take a 128 by 128 tile of a BF16 matrix, 32 KB, and follow it into one SM of an H100.

The tensor memory accelerator copies it from HBM, through the L2, into shared memory. This SM’s fair share of HBM bandwidth is 3.35 TB/s divided by 132, about 25 GB/s, so the copy takes about 1.3 microseconds. From shared memory the tile moves into registers a fragment at a time at 128 bytes per clock, 0.13 microseconds. Multiplying it against another 128 by 128 tile is $2 \cdot 128^3$ FLOPs, 4.2 MFLOP, and this SM’s share of the tensor cores is 7.5 TFLOP/s, so the multiply takes 0.56 microseconds.

The load takes twice as long as the arithmetic it enables. One load, one multiply, and the tensor core is idle more than half the time. Reuse is not an optimisation in this picture. It is the only way the tensor core is ever busy.

Two things make it work. Hopper’s TMA runs the next copy asynchronously while the current tile is being multiplied, so load and compute overlap instead of alternating. And the tile is used far more than once, which is the next question.

Why are tiles 128 wide?

A block of the output, $M_t \times N_t$, needs a strip of $A$ that is $M_t$ rows deep and a strip of $B$ that is $N_t$ columns wide, both running the full length $K$. Once a strip element is in shared memory it is used $N_t$ or $M_t$ times. So the FLOPs per byte that cross the HBM boundary are:

\[I_{\text{tile}} = \frac{2 M_t N_t K}{2K(M_t + N_t)} = \frac{M_t N_t}{M_t + N_t}\]

For a 128-square tile that is 64. For 256, it is 128. Bigger tiles buy intensity linearly, and they are limited by the 227 KB of shared memory and the registers needed to hold the accumulator, which is why 128 and 256 are the numbers you see in every kernel.

Sixty-four FLOPs per byte is still below the SM’s own ridge of about 295. The L2 closes the gap. SMs working on neighbouring output blocks read the same strips of $A$ and $B$, and after the first SM pulls a strip from HBM the rest are served from L2. The effective intensity is the tile’s intensity times the number of SMs sharing each strip, which is why the order in which SMs walk the output matters, and why kernels talk about swizzling and rasterisation.

FlashAttention is this exact picture with $A = Q$, $B = K$ and then $V$, and one more constraint: the output tile never leaves the SM until the whole row of scores has passed through.

Key idea A tile is loaded once and used as many times as it has neighbours. Tile size sets the reuse, shared memory caps the tile size, and the L2 multiplies the reuse across SMs. Every fast kernel is an arrangement of these three facts.

What lives where?

Map a transformer onto the pyramid and the rule is simple: a byte lives at the highest level whose lifetime it fits.

HBM holds everything that must survive between kernels. The weights and their gradients. The optimizer states, which for Adam are the FP32 master weights and two moments. The activations saved for the backward pass. The KV cache during decode. Part 4 counts them.

The L2 holds what several SMs touched a moment ago: the strips of a weight matrix being shared across a matmul, the last few megabytes of an activation on its way to the next kernel.

Shared memory holds the tiles a kernel is working on now. For a matmul, the current $A$ and $B$ tiles. For attention, a block of $Q$ and the blocks of $K$ and $V$ streaming past it, plus a block of scores that never reaches HBM.

Registers, and tensor memory on Blackwell, hold the accumulator and the running statistics: the partial output tile, attention’s running maximum and running sum. Nothing here outlives the kernel.

The roofline’s bytes are the arrows across the HBM boundary. Fusion is the art of moving a computation from the first list to the second.

What changed from Hopper to Blackwell?

Set the two chips side by side and the pattern is stark.

  H100 B200 ratio
dense BF16 tensor cores 989 TFLOP/s 2.25 PFLOP/s 2.3
HBM bandwidth 3.35 TB/s 8 TB/s 2.4
HBM capacity 80 GB 192 GB 2.4
SMs 132 148 1.1
shared memory per SM 227 KB 227 KB 1.0
shared memory bandwidth per SM 128 B/clock 128 B/clock 1.0
exponential units per SM 16 per clock 16 per clock 1.0
tensor memory per SM none 256 KB new
L2 50 MB 126 MB 2.5
NVLink 900 GB/s 1.8 TB/s 2.0

Everything that is a matmul got twice as fast, and HBM kept pace, so the roofline’s ridge barely moved. Everything inside the SM that is not a matmul did not move at all. Shared memory delivers the same bytes per clock. The special function units compute the same sixteen exponentials per clock per SM, against a tensor core that now issues about 8192 FLOPs per clock: a ratio of 512 to one.

Attention needs one exponential per score and $2d$ FLOPs per score. At $d = 128$ the exponential units run dry before the tensor cores do. On Hopper the softmax was a nuisance to hide; on Blackwell it is the bottleneck of the attention kernel, and FlashAttention-4 spends most of its cleverness on it. That is where Part 5 ends.

They will ask Why does FlashAttention not just use a bigger block so the whole row of scores fits in shared memory?
The 30-second version Shared memory is 227 KB per SM. A block of 128 queries against a sequence of 8k keys is a 128 by 8192 score tile, 2 MB in BF16 and 4 MB in FP32, ten to twenty times the budget. And the budget also has to hold the Q block and the current K and V blocks. So the row of scores has to be processed in pieces of a few hundred keys at a time, and the softmax, which needs the whole row's maximum and sum, has to be computed incrementally with a running maximum and a rescaling correction. That correction is the online softmax, and it is the algorithmic idea that makes the hardware constraint survivable.

Rapid fire: can you do these from memory?

  1. List the memory levels of an H100 with sizes and bandwidths, from registers to network.
  2. Roughly how many times faster than HBM is shared memory in aggregate, and where does that number come from?
  3. Time the journey of a 128-square BF16 tile: HBM load, shared-memory transfer, multiply. Which dominates?
  4. Derive the intensity of a tiled matmul at the HBM boundary and evaluate it for 128 and 256.
  5. What does the L2 contribute to a matmul, and why does the SM walk order matter?
  6. Which parts of a training step live in HBM, which in shared memory, which in registers?
  7. Name three things that did not scale from Hopper to Blackwell and one that was added.
  8. Why is the softmax the bottleneck of attention on Blackwell and not on Hopper?

Part 4 goes back to the bottom of the pyramid and counts what fills it: sixteen bytes per parameter, the activation formula, the KV cache, and what sharding does to each.