How do you actually combine them?
Five axes, each derived on its own. In a real job three or four of them run at once, and the question stops being “what does this axis cost” and becomes “which GPUs are in which group”.
That second question has a clean answer, and it is not a set of numbers. It is a matching problem. Rank the collectives by how often they fire and how much they move. Rank the links by bandwidth and latency. Pair them off, busiest to fastest.
Everything else, including the actual degrees, falls out of the memory inventory and the global batch. This part does the matching on the cluster the series has been carrying, then checks the answer against four layouts that have been run at scale and published.
The 30-second version
Start from the byte inventory, not the axes. Then match collectives to links. Tensor parallelism fires four times per layer per micro-batch, so it goes innermost, on NVLink, at the number of GPUs in a node, which for almost everyone is eight. Context parallelism sits next to it if the sequence needs it. Pipeline parallelism goes across nodes, deep enough that the model states fit and no deeper, because every stage costs bubble. Fully sharded data parallelism goes outermost, because its collective fires once per optimizer step and can be prefetched and overlapped across a multi-hop network. Meta wrote that order down for Llama 3 as tensor, context, pipeline, data, innermost first, and gave exactly that reasoning. Then set the micro-batch count as large as the global batch allows, since the bubble is the pipeline depth minus one over it. Check the answer against published layouts: Llama 3 405B ran tensor 8, pipeline 16, data 64 to 128 on eight to sixteen thousand H100s at 38 to 43% MFU. Megatron's trillion-parameter run was tensor 8, pipeline 64, data 6 at 52% of peak. DeepSeek-V3 used no tensor parallelism at all, because expert parallelism was already using the node's bandwidth.What does the machine actually look like?
Before choosing any degree, draw it. 512 H100s in 64 nodes of eight.
Inside a node, every GPU reaches every other over NVLink at 450 GB/s per direction. Between nodes, each GPU has a 400 Gb/s port, which is 50 GB/s. That is a factor of nine, and it is the only fact about the hardware that the layout has to respect.
At real scale there is a third tier. Meta describe the Llama 3 cluster as racks of 16 GPUs across two servers under a top-of-rack switch, 192 racks forming a 3072-GPU pod with full bisection bandwidth, and eight pods forming a 24,000-GPU cluster with a 1:7 oversubscription ratio above the pods. Their scheduler and their parallelism layout are both topology-aware, and the goal is the same in both: keep the heaviest traffic off the oversubscribed tier.
So the design rule is one sentence. A collective that runs four times per layer per micro-batch must stay on the fast link. A collective that runs once per optimizer step and can hide behind the backward pass may use the slow one.
Which axis gets which link?
Give the GPUs global indices and cut them into groups. Which axis owns the consecutive indices is the decision, because consecutive indices are what share a node.
With the tensor degree innermost, a tensor-parallel group is a run of $t$ consecutive GPUs. At $t = 8$ each group is exactly one node, and all 320 of its collectives per micro-batch stay on NVLink. The second step of the figure colours the cluster by group, and every node is a single colour.
Push $t$ to 16 and every group straddles two nodes. The figure outlines every node in red, and the arithmetic underneath moves: the same 80.5 GB per micro-batch takes 1.61 seconds on the network instead of 179 milliseconds on NVLink, against 219 milliseconds of arithmetic. One slider notch turns a compute-bound step into a network-bound one.
That is the whole argument, and it explains why the tensor degree in every published dense recipe is eight. It is a property of the machine, not of the model.
What order do the dimensions go in?
Meta state the rule for Llama 3 explicitly: the dimensions are ordered tensor, context, pipeline, data, innermost first, because the innermost needs the highest bandwidth and lowest latency and the outermost can tolerate a multi-hop network.
Their reason for putting data parallelism outermost is worth quoting in spirit: fully sharded data parallelism prefetches its weight all-gathers and reduces its gradients asynchronously, so it tolerates latency in a way that a synchronous all-reduce in the middle of a block cannot.
Megatron’s guidance from 2021 is the same conclusion reached from the other side, and it comes in two takeaways. Use tensor parallelism up to the number of GPUs in a server, then use pipeline parallelism to go further. And choose the total model-parallel size, tensor times pipeline, so that the parameters and their states fit, then scale out with data parallelism.
What does a layout cost, term by term?
The third step of the figure puts a clock on it, with a deliberately simple model: arithmetic at peak, tensor-parallel collectives fully exposed, a bubble of $(p-1)/m$, point-to-point sends between stages, and fully sharded data-parallel traffic hidden behind the backward pass except for whatever sticks out.
It is not a simulator, and the kernels’ own distance from the roofline is not in it. What it shows is where each axis puts its cost.
At tensor 8, no pipeline, data 64, and 32 micro-batches, the step is 19.4 seconds of which 72% is arithmetic and 28% is tensor-parallel collectives. Multiply that useful fraction by whatever the kernels themselves achieve against the roofline and you get something in the neighbourhood of a real MFU number.
Three things fall out of dragging the sliders. Tensor 16 moves the collectives onto the network and the red block swallows the step. A deep pipeline with few micro-batches makes the bubble the dominant term. And a small micro-batch count shrinks the compute that the data-parallel traffic has to hide behind, which is Part 2’s third wall arriving in a new place.
What have people actually run?
Four published layouts, and they are the fastest way to sanity-check your own.
Meta’s Llama 3 405B ran three configurations during pre-training. On 8192 H100s: tensor 8, context 1, pipeline 16, data 64, 8192-token sequences, 16 million tokens per batch, 430 TFLOP/s per GPU and 43% BF16 MFU. On 16,384 H100s: the same except data 128, and 400 TFLOP/s at 41%. And for the long-context stage on the same 16,384 GPUs: tensor 8, context 16, pipeline 16, data 8, on 131,072-token sequences, 380 TFLOP/s at 38%.
Megatron-LM’s trillion-parameter run was tensor 8, pipeline 64, data 6 on 3072 A100s, achieving 163 TFLOP/s per GPU, 52% of peak, and 502 PFLOP/s aggregate.
DeepSeek-V3 is the interesting one because it breaks the pattern. On 2048 H800s: no tensor parallelism at all, 16-way pipeline with their DualPipe schedule, 64-way expert parallelism spanning eight nodes, and ZeRO-1 data parallelism over the rest. They say plainly why: careful memory engineering let them avoid the cost of tensor parallelism, and the node’s bandwidth was already committed to the expert all-to-all.
What happens to efficiency as you add an axis?
It falls, every time, and the published numbers put a size on it.
Llama 3 405B lost two points of MFU, 43% to 41%, purely from doubling the data-parallel degree at a fixed global batch, because the per-rank batch halved. It lost three more, to 38%, when context parallelism came in for the long-context stage. Nothing else in the configuration changed.
Meta’s PyTorch team report the same shape from the other direction with torchtitan: stacking optimisations gave a 65% speedup on Llama 3.1 8B at 128 GPUs with one dimension of parallelism, 12.6% on 70B at 256 GPUs with two, and 30% on 405B at 512 GPUs with three, all against already-optimised baselines. The absolute numbers are not comparable across rows, but the direction is: each dimension you add is harder to make efficient than the one before it.
So the order of operations is not “add axes until it fits”. It is “add the fewest axes that make it fit, in the order the topology dictates”.
How much of the communication can actually be hidden?
Different amounts, and knowing which is which is the difference between a good answer and a list of technique names.
Fully sharded data parallelism hides almost everything, if the wrapping is right. The all-gather for the next unit is issued while the current unit computes, so only the very first gather in the model is exposed.
Pipeline point-to-point hides well, because it is small and asynchronous, and Meta report that making it asynchronous mattered most when a document mask made stages imbalanced.
Tensor-parallel collectives hide badly, because they sit between two dependent GEMMs. Asynchronous tensor parallelism is the current answer: decompose the matmul and the collective into chunks so one chunk’s transfer overlaps the next chunk’s arithmetic.
The all-to-all of expert parallelism hides worst of all, which is why DeepSeek built an entire bidirectional pipeline schedule to overlap one micro-batch’s communication with another’s computation, and hand-wrote kernels that use twenty streaming multiprocessors so the rest of the chip keeps computing.
There is also a layer below all of this. Meta fork NCCL for the same reason: at tens of microseconds of latency, the chunking and staging inside the collective library itself become the bottleneck, and they had to tune it and give small control messages priority so they do not get stuck behind bulk traffic in a deep-buffer switch.
What changes when the run has to survive failures?
The layout decides the blast radius of a dead GPU, and that never appears in a throughput table.
A tensor-parallel group is a single logical device: lose one member and the whole group is down. A pipeline is a chain: lose one stage and every stage stops. A context-parallel ring is the same. An expert-parallel group loses the experts hosted on that rank, so the layer cannot run.
A data-parallel replica is the only genuinely redundant axis, and even then only if the framework can re-form the group without the dead rank and rescale the batch.
So deep pipelines and wide tensor groups both enlarge the set of GPUs that must be alive at the same instant, which at a fixed cluster size shortens the mean time between restarts. The failure rates themselves, and the checkpoint interval that follows from them, are in Part 7 of the GPU series.
The practical consequence for the layout is checkpoint format. The next run may have different degrees, either because a node died or because you changed your mind, so the checkpoint has to store logical tensors and reshard on load rather than dumping per-rank buffers. Every framework’s distributed checkpoint works that way, and it is the reason it exists.
How would you choose, from scratch?
The order I would say out loud, with the reason attached to each step.
Write the byte inventory first: sixteen bytes per parameter, $34\,s\,b\,h$ per layer, plus overheads. Nothing else is decidable until that number exists.
Take tensor parallelism up to the GPUs in one node, because it is the only axis that divides weights and activations together, and it needs NVLink. Stop at the key-value head count if that is smaller.
Add context parallelism only if the sequence is what does not fit, and put it next to the tensor dimension.
Add pipeline stages only if the model states still do not fit, deep enough and no deeper, because each stage costs bubble and blast radius.
Fill the rest with fully sharded data parallelism, which is the only axis that scales with the cluster.
Then set the micro-batch count as large as the global batch allows, and if the bubble still hurts, interleave or use a zero-bubble schedule before touching any of the degrees above.
The 30-second version
It trades a real cost for a different real cost, so it depends on what is binding. Tensor 4 halves the collective traffic per micro-batch relative to tensor 8, since the ring factor barely changes but the layers per group do not, and it keeps the GEMMs twice as wide, which helps the kernels. What it gives up is activation memory: tensor parallelism with sequence parallelism divides activations by the degree, so tensor 4 leaves twice as much per GPU, and the pipeline stage inside the node does not give it back, because a stage still holds the whole model's worth of activations for the micro-batches in flight. It also puts a pipeline boundary on NVLink, which is a waste of the fastest link in the machine on the cheapest collective in the job. I would take it seriously only when the activation budget is comfortable and the tensor collectives are measurably the bottleneck, and I would measure rather than argue, because both effects are within a factor of two of each other.Rapid fire: can you do these from memory?
- State the ordering of parallelism dimensions and give the reason for each end of it.
- Why is the tensor degree eight in almost every published dense recipe?
- Give the Llama 3 405B configuration for the 8k stage and for the long-context stage, and the MFU of each.
- Name the four terms of a step-time budget for a 3D layout.
- Rank the four main collectives by how well they hide, and say why.
- Which axes are brittle to a single GPU failure, and which is elastic?
- Give the order you would choose the degrees in, and the reason attached to each step.
Part 7 is the question set: three levels, and no figures, just whether it stuck.