Is this kernel memory-bound or compute-bound?
Part 1 kept dividing FLOPs by bytes. That ratio has a name, arithmetic intensity, and a chart, the roofline, and together they answer the question every performance conversation eventually reaches: is this thing waiting on the tensor cores or on HBM?
The roofline was drawn for CPUs in 2009 by Williams, Waterman and Patterson. It has become the one picture GPU people agree on, because it takes two numbers about the machine and one number about the kernel and tells you the speed limit. Interviewers like it because it separates people who have measured from people who have read.
I want to put every kernel in a transformer on it. Some of the placements are surprising.
The 30-second version
Arithmetic intensity is FLOPs performed per byte moved from HBM. A kernel cannot run faster than the peak FLOP/s, and it cannot run faster than bandwidth times intensity, because each byte buys at most that many FLOPs. Attainable performance is the minimum of the two, which drawn on log axes is a slope meeting a flat roof. The corner is the ridge point, peak over bandwidth: about 295 FLOP per byte for an H100 in BF16, 281 for a B200. Left of it a kernel is memory-bound and only bandwidth helps; right of it, compute-bound. Big matmuls sit far right. Norms, softmax, residual adds and single-token decode sit at one to three FLOP per byte, far left. Naive attention materialises the score matrix and sits at intensity d/2, about 64, memory-bound at every sequence length. Fused attention keeps the scores on chip, moves only Q, K, V and O, and has intensity N/2, so it is compute-bound past a few hundred tokens. Same FLOPs, a hundred times fewer bytes.What are the two ceilings?
For a kernel, count two things. The FLOPs it performs, and the bytes it moves in and out of HBM. Their ratio is the arithmetic intensity:
\[I = \frac{\text{FLOPs}}{\text{bytes moved}}\]The machine caps the kernel twice. It cannot exceed the tensor cores’ peak, a flat line. And every byte it needs arrives at the HBM bandwidth and buys at most $I$ FLOPs, so it cannot exceed bandwidth times intensity, a diagonal. What it can attain is the lower of the two:
\[\text{attainable FLOP/s} = \min\!\big(\text{peak},\ \text{BW} \cdot I\big)\]On log axes that is a slope that bends into a roof. Every kernel is a point under it. Step through the figure; it puts the whole transformer on the chart, one kernel at a time.
Where is the ridge point?
The corner is where the two ceilings meet, at $I^* = \text{peak} / \text{BW}$. It is the machine’s balance: the FLOPs it can do in the time one byte arrives.
| machine | peak, dense BF16 | HBM bandwidth | ridge $I^*$ |
|---|---|---|---|
| A100 SXM | 312 TFLOP/s | 2.04 TB/s | 153 |
| H100 SXM | 989 TFLOP/s | 3.35 TB/s | 295 |
| B200 | 2.25 PFLOP/s | 8 TB/s | 281 |
Left of the ridge a kernel is memory-bound: more bandwidth speeds it up, more FLOP/s does nothing. Right of it, compute-bound. The roof rose seven times in five years and the ridge moved by two, because compute and bandwidth grew together. Where a kernel sits has barely changed. How fast it runs has.
Lower precision breaks that pattern. In FP8 the B200’s peak is 4.5 PFLOP/s and in FP4 it is 9, with the same 8 TB/s underneath, so the ridge doubles each step: 281, 562, 1125. More on that at the end.
Why is a big matmul compute-bound and a skinny one not?
A square matmul of size $N$ performs $2N^3$ FLOPs on three $N \times N$ operands, so at two bytes per element its intensity is $N/3$. It walks right as it grows and crosses the H100’s ridge near $N = 885$. Anything larger is compute-bound.
Transformer matmuls are rarely square. Streaming $T$ tokens through the FFN weight, $T \times 8192$ by $8192 \times 28672$, moves the weight once and the activations once. For small $T$ the weight dominates the bytes, so the intensity is about $T$: one token buys one FLOP per byte of weight.
That is the entire theory of batching in one sentence. Tokens per weight read is the intensity, and the ridge point is a batch size. Below about 300 tokens per weight read on an H100 the tensor cores idle; above it, HBM does.
Why can decode never reach the roof at long context?
Decode with a batch of $B$ sequences reads every weight once, which is the skinny matmul with $I \approx B$. So batch a few hundred sequences and you should reach the ridge.
You do not, and the reason is the KV cache. Every sequence reads its whole cache each step, 320 KiB per token of context for Llama 3 70B, and that traffic grows with the batch exactly as fast as the FLOPs do. Take the limit:
\[I \;\to\; \frac{2P + 4 L h L_{\text{layers}}}{L \cdot 320\,\text{KB}} \quad \text{as } B \to \infty\]where $P$ is the parameter count, $L$ the context and the second term in the numerator is the attention arithmetic. At 8k context the cap is about 60 FLOP per byte. At 128k it is about 11. Both are far left of every ridge.
So decode at long context is memory-bound at every batch size, and no amount of batching fixes it. The only way up is fewer cache bytes per FLOP, which is what grouped-query attention, multi-head latent attention, FP8 caches and sparse attention all do. Part 6 collects them.
Which kernels live on the slope?
Most of them. A residual add reads two tensors and writes one: six bytes per FLOP. RMSNorm reads and writes every element once for about six FLOPs per element. Softmax and SwiGLU have the same shape. All of them sit between 0.2 and 3 FLOP per byte, a hundred times left of the ridge.
On an H100 a kernel at intensity 1.5 attains 5 TFLOP/s, half a percent of peak, with HBM saturated the whole time. Its duration is bytes divided by bandwidth and nothing else. A norm over an $8192 \times 8192$ BF16 activation reads 134 MB and writes 134 MB: 80 microseconds, during which its 0.4 GFLOP would have taken 400 nanoseconds.
A layer has a dozen such kernels, and together they cost about as much as one of its matmuls. The only fix is fewer bytes. Fuse the norm into the matmul that reads its output. Compute the activation function in the matmul’s epilogue before the result is written. Never write the intermediate at all.
What does the roofline say about attention?
Attention over $N$ tokens with head width $d$ performs $4N^2 d$ FLOPs. Written the textbook way, it writes the $N \times N$ score matrix to HBM, reads it back for the softmax, writes the probabilities, and reads them again to multiply by $V$. That is $8N^2$ bytes, so:
\[I_{\text{naive}} = \frac{4N^2 d}{8N^2} = \frac{d}{2} = 64\]Independent of $N$, and left of every ridge. Naive attention is memory-bound at any sequence length, which is why it was slow long before it ran out of memory.
Keep the scores on chip and only $Q$, $K$, $V$ and the output touch HBM: $8Nd$ bytes.
\[I_{\text{fused}} = \frac{4N^2 d}{8Nd} = \frac{N}{2}\]At 8k tokens that is 4096, far right of the ridge. Same FLOPs, a hundred times fewer bytes, and the kernel jumps from the slope to the roof. The roofline predicted FlashAttention before anyone wrote it. How the scores stay on chip when the matrix is 128 MB is Part 3’s question, and Part 5 writes the algorithm.
Does lower precision help?
It moves the roof, not the slope. On a B200, BF16 peaks at 2.25 PFLOP/s, FP8 at 4.5 and FP4 at 9, all over the same 8 TB/s. Each halving of precision doubles the ridge.
The consequence cuts both ways. A kernel whose bytes stay in BF16 becomes more memory-bound every time the tensor cores get faster; its point does not move while the ridge runs away from it. A kernel whose bytes shrink with its arithmetic slides right exactly as far as the ridge does, so the crossover batch stays near 281 tokens in every precision and the time halves each step.
There is a second asymmetry, and it is the defining fact of Blackwell. Between the H100 and the B200 the tensor cores got 2.25 times faster, but shared-memory bandwidth per SM and the number of exponential units did not change. Everything that is not a matmul got relatively slower, including the softmax inside attention. FlashAttention-4 exists because of this, and Part 5 ends there.
How do I count the bytes for a kernel?
The recipe is short, and it is what I do before touching a profiler.
FLOPs: for a matmul, $2MNK$. For elementwise kernels, a handful per element. For attention, $4N^2 d$ per head.
Bytes: every input read and every output written, at its precision, counted once per pass over HBM. A matmul moves its three operands. A fused kernel moves less than the sum of its parts, which is the point of fusing.
Then $I$ is the ratio, and the time is the larger of FLOPs over peak and bytes over bandwidth. Ten lines of Python do it:
def kernel_time(flops, bytes_moved, peak_flops=989e12, bw=3.35e12):
intensity = flops / bytes_moved
ridge = peak_flops / bw
t = max(flops / peak_flops, bytes_moved / bw)
bound = "compute" if intensity > ridge else "memory"
return intensity, t, bound
# RMSNorm over an 8192 x 8192 bf16 activation
print(kernel_time(flops=6 * 8192 * 8192, bytes_moved=2 * 2 * 8192 * 8192))
# (1.5, 8.0e-05, 'memory')
# the FFN up-projection with 8192 tokens
T, h, f = 8192, 8192, 28672
print(kernel_time(flops=2 * T * h * f, bytes_moved=2 * (T * h + h * f + T * f)))
# (3591.1, 3.9e-03, 'compute')
Measured numbers land under the roof, never on it: a good matmul reaches 70 to 80% of peak, a good memory-bound kernel 80 to 90% of bandwidth. The gap between the roof and the point is what kernel engineers are paid to close.
The 30-second version
Not necessarily. Compute its intensity first. If it moves a byte per FLOP, the roofline says 5% of peak is about the most an H100 can give it, and it is running at bandwidth. The question is then whether the bytes are necessary, and the fix is fusion, lower precision on the bytes, or restructuring so intermediates stay on chip, not a faster inner loop. If its intensity is in the thousands and it still runs at 5%, then yes, the kernel is leaving the tensor cores idle: bad tiling, poor occupancy, or waiting on memory it could have prefetched.Rapid fire: can you do these from memory?
- Define arithmetic intensity and write the attainable FLOP/s as a minimum of two terms.
- Derive the ridge point and give it for an H100 in BF16 and a B200 in FP8.
- Why is the intensity of a square matmul N/3 and of a skinny one about the number of tokens?
- Show that decode intensity saturates as the batch grows, and at what value for 8k context.
- Where do RMSNorm and the residual add sit, and why does fusion help them?
- Derive the intensity of naive and fused attention. Which side of the ridge is each?
- Does moving to FP8 make a BF16-byte kernel more or less memory-bound? Why?
- Count the bytes and FLOPs for a 4096 by 4096 by 4096 matmul in BF16 and place it.
The roofline counts bytes to and from HBM as if HBM were the only memory. It is not. Part 3 opens the chip: registers, shared memory, tensor memory, the L2, and the thousand-fold spread in bandwidth between them that makes fused kernels possible.