What is state of the art in September 2026?

The first five parts built a vocabulary: four clocks, a roofline, a pyramid, an inventory of bytes, and one kernel that used all of them. This part is the catalogue. Everything below is in production as I write, in the default paths of PyTorch, Megatron, vLLM, SGLang and TensorRT-LLM, or in the training reports of the frontier labs this year.

I have organised it the only way that makes it memorable: by the bottleneck each technique relieves. Every chip on the map in the figure sits on one of Part 1’s lines. If you can say which line a technique moves, you understand it. If you can also say the number it moves it by, you are ready for the follow-up.

The dates matter. Half of this did not exist two years ago, and an interviewer in 2026 will notice a 2024 answer.

They will ask What are the most important optimisations in large-model training and serving today, and what does each one buy?
The 30-second version Group them by bottleneck. Against compute: low precision, FP8 with fine-grained scales as DeepSeek-V3 did, and now NVFP4 pretraining on Blackwell, which doubles the tensor-core roof each step and trained Llama 3.1 405B 3.2 times faster than Hopper FP8 in MLPerf; plus compilers and kernel DSLs that reach the roof. Against HBM bandwidth: FlashAttention, kernel fusion, and for serving the KV cache shrinkers, grouped-query and multi-head latent attention, FP8 caches, DeepSeek's sparse attention, and speculative decoding. Against capacity: FSDP2 and ZeRO, tensor, sequence and context parallelism, recompute, and cheaper optimizers like Muon and 8-bit Adam. Against the units inside an SM: FlashAttention-3 and 4. Against the interconnect: zero-bubble and bidirectional pipelines, overlapped sharding, NVL72 racks and 800 gigabit fabrics, and disaggregated prefill and decode with RDMA cache transfer. Against launch overhead: CUDA graphs and torch.compile. Against failures: asynchronous checkpointing and fault-tolerant schedulers.

Which technique attacks which bottleneck?

Seven rows, because there are seven things a job can run out of: tensor-core throughput, HBM bandwidth, HBM capacity, the units inside an SM, the links between GPUs, the host that launches the kernels, and the hardware itself. Two columns, training and serving. The figure’s first step is the map; the rest of it takes the rows one at a time.

How low has precision gone?

Precision is the one lever that raises the roof and lowers the bytes at the same time, so it helps compute-bound and memory-bound kernels alike, and it is where the largest gains of the last two years came from.

DeepSeek-V3, in December 2024, was the first open model to validate an FP8 mixed-precision framework at frontier scale. Mixed is the operative word: the expensive GEMMs run in FP8, while the embedding, the output head, the MoE gating, the normalisations and the attention operators are kept at BF16 or FP32, and master weights, gradients and optimizer states are held higher. The trick that made the FP8 half stable was fine-grained scaling: one scale per 1 by 128 tile of activations and per 128 by 128 block of weights, so a single outlier cannot flatten a whole matrix, with accumulation promoted to FP32 every few tiles. Loss within 0.25% of BF16.

Blackwell hardened the idea into formats. MXFP8 gives every 32 values a shared power-of-two scale. NVFP4 gives every 16 values an FP8 scale and the whole tensor one FP32 scale, four bits per value. NVIDIA’s pretraining recipe for it, published in September 2025 and refined since, has four ingredients: random Hadamard rotations before quantising gradients, so outliers spread across a block; stochastic rounding, so the quantisation error has zero mean; two-dimensional scaling for weights; and the most sensitive layers left in higher precision. It trained a 12B hybrid Mamba-Transformer on ten trillion tokens within 1.5% of the FP8 loss.

In MLPerf Training v5.1, NVFP4 on a GB200 NVL72 rack trained Llama 3.1 405B 3.2 times faster than Hopper in FP8 at the same GPU count. By July 2026, MXFP8 and NVFP4 run reinforcement learning end to end, rollouts and updates in the same format. The roofline caveat from Part 2 stands: each halving doubles the ridge, so a kernel only benefits if its bytes shrink with it, and activations and caches have lagged the weights.

What happened to attention and the cache?

Part 2 found the KV cache capping decode intensity far below the ridge, and Part 4 found it capping how many sequences a GPU can hold. The fixes come in three families, and a strong answer names one from each.

Store less per token. Grouped-query attention cut Llama’s cache eight-fold, 320 KiB per token instead of 2.5 MiB. Multi-head latent attention, in DeepSeek-V2 and V3, compresses keys and values into a 576-wide latent per layer, about 69 KiB per token across 61 layers. FP8 caches halve whichever you have.

Read less of it. DeepSeek Sparse Attention shipped in V3.2 in December 2025. A lightweight lightning indexer scores every earlier token against the current query with a few ReLU-gated dot products in FP8, each query attends only to its top 2048 keys, and the model was continued from V3.1 in two stages rather than retrained. At 128k context the cost falls three to six times with quality matched to the dense model. Native Sparse Attention, from the same lab in early 2025, is the trained-from-scratch cousin.

Read it for more than one token. Speculative decoding drafts several tokens and verifies them in one forward pass, so each read of the weights and cache produces two to four tokens instead of one. EAGLE-3, which drafts from the target model’s own hidden states, is the production standard in vLLM, SGLang and TensorRT-LLM. P-EAGLE, from March 2026, drafts all of its tokens in a single pass of the drafter and adds 1.05 to 1.69 times on top of EAGLE-3 on B200s.

All three families move the same number, FLOPs per byte of cache read, and together they are why a 128k-context model is a product in 2026 and not a demonstration.

Key idea Name the line a technique moves and the factor it moves it by. GQA divides cache bytes by eight, MLA by another four, sparse attention divides reads by sixty-four at 128k, speculative decoding multiplies tokens per read by three. That sentence is the interview answer.

How do the frameworks shard and schedule now?

The training frameworks converged on a shape. PyTorch’s FSDP2, the recommended path since 2.6, shards weights, gradients and optimizer states per module through a DTensor-based API, and prefetches each layer’s all-gather a layer ahead so the communication hides under compute. Megatron shipped its own FSDP in 2026 with pooled allocation and the overlap tuned for its tensor and pipeline groups. The ZeRO arithmetic from Part 4 is what both implement.

Pipeline parallelism lost most of its idle time. The zero-bubble schedules split each micro-batch’s backward pass into its two halves, the gradient with respect to activations, which the next stage is waiting for, and the gradient with respect to weights, which nobody is waiting for, and slot the weight halves into what were bubbles. DeepSeek-V3’s DualPipe runs the pipeline in both directions at once, so the all-to-all traffic of its 64-way expert parallelism hides under the compute travelling the other way.

Two recipes to know by heart. DeepSeek-V3: 16-way pipeline, 64-way expert parallelism across 8 nodes, ZeRO-1 data parallelism, 671B parameters on 2048 H800s. Llama 3 405B: tensor 8 inside the node, pipeline 16, context 16 for the long-context stage, FSDP across the remaining data-parallel dimension, 16,384 H100s. Each number is a division of a line in Part 4’s inventory.

Did the optimizer change?

For most models, no: AdamW is still the default. But two of the sixteen bytes per parameter came under attack, and one of the attacks trained a frontier model.

Muon replaces Adam’s two moments with a single momentum buffer and orthogonalises each weight matrix’s update with a few Newton-Schulz iterations. Moonshot scaled it to a trillion-parameter mixture of experts in Kimi K2 in July 2025: 15.5 trillion tokens, and, they report, no loss spikes over the run and about twice AdamW’s efficiency per FLOP. The stability came from a second idea, QK-Clip, which rescales a head’s query and key projections after any step in which its attention logits grew past a threshold. It is the training-time cousin of the QK-norm from the Transformer series, applied to the weights after the fact.

8-bit Adam keeps both moments in blockwise-quantised INT8, cutting the optimizer state from twelve bytes per parameter to about six. Muon cuts it to eight. At 70B parameters those are 280 and 420 GB of HBM across the cluster, which is either more GPUs freed or longer sequences fitted.

What does serving look like now?

Serving has two phases with opposite bottlenecks, and 2026 is the year the industry stopped running them on the same GPU.

Prefill processes the whole prompt at once, thousands of tokens per weight read: compute-bound, on the roof. Decode produces one token per sequence per step: memory-bound, on the slope, with a user waiting on every millisecond. On a shared GPU a long prefill stalls every decoding user behind it, and the two phases want different batch sizes, different parallelism and, increasingly, different hardware.

Disaggregation gives each its own pool. Prefill GPUs build the KV cache and hand it to decode GPUs, which stream tokens. The handover is the catch: 2.7 GB for an 8k prompt of Llama 3 70B. NVIDIA’s NIXL library moves the cache tensors directly between GPU memories over RDMA, and in 2026 it is the standard path in both vLLM and Dynamo, with llm-d doing the same inside Kubernetes. DistServe and Splitwise, the 2024 papers, measured up to 4.5 times more requests within the same latency target. It is not free: below a few sequences per GPU the transfer costs more than it saves, and it needs an RDMA fabric.

Underneath sit the defaults every engine ships. Continuous batching admits new sequences between decode steps instead of waiting for the batch to drain. Paged attention keeps the cache in fixed blocks with a page table, so no sequence reserves its maximum length. Prefix caching computes a shared system prompt once. CUDA graphs run the decode step, and a KV-aware router sends each request to the GPU that already holds its prefix.

What hardware shipped this year?

chip year memory bandwidth dense BF16 dense FP8 dense FP4 BF16 ridge
H100 SXM 2022 80 GB HBM3 3.35 TB/s 989 TF 1.98 PF   295
H200 2024 141 GB HBM3e 4.8 TB/s 989 TF 1.98 PF   206
B200 2025 192 GB HBM3e 8 TB/s 2.25 PF 4.5 PF 9 PF 281
B300, Blackwell Ultra 2025 288 GB HBM3e 8 TB/s     15 PF  
Rubin, VR200 H2 2026 288 GB HBM4 up to 22 TB/s     35 PF train, 50 PF infer  
AMD MI355X 2025 288 GB HBM3e 8 TB/s 2.5 PF 5 PF 10 PF 313
AMD MI400 2026, announced 432 GB HBM4 19.6 TB/s        
Google TPU v7 Ironwood GA April 2026 192 GiB HBM3e 7.4 TB/s   4.6 PF    
AWS Trainium 3 2025 to 2026 144 GB HBM3e 4.9 TB/s        

All figures dense, without structured sparsity. Between racks, Blackwell systems use ConnectX-8 at 800 gigabits per GPU into Quantum-X800 InfiniBand or Spectrum-X Ethernet; inside the rack, NVLink 5 at 1.8 TB/s per GPU links 72 GPUs with 130 TB/s of aggregate bandwidth. Rubin’s rack was announced as NVL144 by die count and ships as VR200 NVL72 by package. The MI400 figures are AMD’s announcement and not yet independently measured.

Read the last column. Four years, and the BF16 ridge point barely moved. Memory grew about two and a half times. The FP4 roof is where a decade of gains landed, and every technique on this map exists to reach that roof with bytes that did not shrink as fast.

They will ask You are moving a training run from H100s to B200s. What changes, and what do you have to re-examine?
The 30-second version Compute and bandwidth both grew about 2.4 times, so a BF16 job that was compute-bound stays compute-bound and runs about twice as fast with no changes. Memory grew 2.4 times too, so the sharding can relax: fewer GPUs per model replica, larger micro-batches, less recompute, which raises arithmetic intensity further. The real gains need work. Switching to FP8 or NVFP4 doubles or quadruples the roof, but only for kernels whose bytes shrink with it, so re-examine which activations and caches are still in BF16, and keep a BF16 control run to compare loss curves. Anything that is not a matmul got relatively slower: check the softmax-heavy kernels are on FlashAttention-4 rather than Triton, and check the norms and elementwise ops are fused. Interconnect doubled to 1.8 TB/s NVLink and 800 gigabit ports, and the NVL72 domain is 72 GPUs, so the parallelism layout that assumed an 8-GPU NVLink island should be redrawn.

Rapid fire: can you do these from memory?

  1. Name one technique per bottleneck row, for training and for serving.
  2. What made DeepSeek-V3's FP8 training stable, and what are the four ingredients of the NVFP4 recipe?
  3. Give the KV bytes per token for MHA, GQA and MLA on a Llama-70B-shaped model.
  4. How does DeepSeek Sparse Attention choose what to attend to, and what does it cost at 128k?
  5. Why does speculative decoding help a memory-bound decode step, and what is P-EAGLE's change over EAGLE-3?
  6. What does a zero-bubble pipeline schedule move into the bubbles, and why is that legal?
  7. How many bytes of optimizer state per parameter for Adam, 8-bit Adam and Muon?
  8. Why disaggregate prefill and decode, what is the cost, and what made it routine?

Part 7 is what none of these papers put in the abstract: what actually goes wrong when you run at scale, from the 40% of peak that the best runs achieve to the GPU that fails every three hours, and the short list of things to check before anything else.