How do you split the layers?

Tensor parallelism cut across every layer and paid in bandwidth. The other way to cut a model is between layers: give the first twenty to one group of GPUs, the next twenty to the next, and pass activations down the line.

It is the cheapest axis on the wire by a very wide margin. One activation tensor per stage boundary per micro-batch, point to point, no collective at all. It is also the only axis whose cost is idle time, and idle time does not overlap with anything.

Everything about pipeline parallelism is a fight with that idle time. There are four schedules worth being able to draw, and the difference between them is entirely in when the backward passes run.

They will ask What is the pipeline bubble, how big is it, and how do you make it smaller?
The 30-second version The bubble is the idle time at the start and end of every batch, while the pipeline fills and drains. It is p minus one slots of warm-up and the same again of drain, against m micro-batches of useful work, so the bubble is (p−1)/m of the ideal time and you want m much larger than p. GPipe runs all forwards then all backwards and holds all m micro-batches' activations; 1F1B alternates once warmed up, has exactly the same wall clock, and caps in-flight micro-batches at p minus the stage index, which is what makes large m affordable. Interleaving gives each device v non-contiguous chunks of the model so the slot is v times shorter, cutting the bubble to (p−1)/(vm) at the cost of v times as many point-to-point messages. Zero-bubble schedules split the backward pass into the input gradient, which the previous stage is waiting for, and the weight gradient, which nobody is waiting for, and use the second as filler; that takes the bubble to a third with no extra memory, or to nothing at all with about twice the activations in flight. And the memory floor that none of them fix: the first stage holds p micro-batches of L over p layers each, which is L layers' worth of activations whatever p is.

What crosses a stage boundary?

Almost nothing, which is the point.

The forward pass sends one activation tensor from the last layer of stage $i$ to the first layer of stage $i+1$. With tensor and sequence parallelism already applied, that tensor is $s/t \times h$ in BF16, which for our shape is 16.8 MB. The backward pass sends the same shape the other way.

Compare that with the 75.2 GB per micro-batch that tensor parallelism moves in the same pass. Pipeline traffic is three orders of magnitude smaller, and it is point to point rather than a collective, so it does not synchronise anyone who is not directly involved.

The weights, gradients and optimizer states all divide cleanly by $p$. At four stages our 70B model’s states go from 141 GB per tensor-parallel rank to 35.3 GB.

And then you run one batch through it and every stage but one is idle.

Where does the bubble come from?

Cut the batch into $m$ micro-batches and push them in one after another. Stage 0 starts on micro-batch 1 at time zero, stage 1 starts on it one slot later, and stage $p-1$ cannot start until $p-1$ slots have passed.

The same thing happens in reverse at the end, when the last micro-batch drains back down the pipeline. Between those two, every stage is busy.

\[t_{\text{bubble}} = (p-1)(t_F + t_B), \qquad t_{\text{ideal}} = m(t_F + t_B), \qquad \frac{t_{\text{bubble}}}{t_{\text{ideal}}} = \frac{p-1}{m}\]

At $p = 4$ and $m = 16$ that is 18.8%, and the figure measures exactly that off the drawn schedule rather than asserting it. Drag $m$ and the number falls as $1/m$. Drag $p$ and it climbs.

A flush at the end of every batch is what forces this. Asynchronous and bounded-staleness pipelines avoid it by letting different micro-batches see different weight versions, and every production framework rejected that trade, because reproducing a training curve matters more than the last few percent.

Why does 1F1B not make it faster?

Because the bubble does not care what order the passes run in. It cares how deep the pipeline is.

In GPipe, each stage runs all $m$ forward passes, then all $m$ backward passes. In 1F1B, a stage warms up with a few forwards and then alternates one forward and one backward for the rest of the batch. Put the two schedules next to each other on the same axis, as the third step of the figure does, and the wall clocks are identical.

What changes is the number of micro-batches a stage has in flight, meaning micro-batches whose forward pass has run and whose backward pass has not, so whose activations must still be resident. GPipe holds all $m$. 1F1B holds at most $p - i$ on stage $i$.

At $m = 16$ and four stages, that is 91.3 GB of activations on stage 0 against 22.8 GB. The first number does not fit in an H100 and the second does.

So 1F1B does not make the pipeline faster. It makes a large $m$ affordable, and a large $m$ is what makes the pipeline fast. Megatron calls it PipeDream-Flush and it has been the default everywhere for years.

Key idea 1F1B buys memory, not time. Interleaving and zero-bubble scheduling buy time. Knowing which schedule bought which is the whole of this subject.

How does interleaving cut the bubble?

The bubble is $p-1$ slots long, and a slot is however long one stage’s forward pass takes. So make the slot shorter.

Give each device $v$ non-contiguous chunks of the model instead of one contiguous block. With $v = 2$ and eight stages, device 0 owns the first sixteenth of the layers and the ninth sixteenth. The pipeline now has $vp$ virtual stages, each a $v$-th as long.

The last device still starts after $p-1$ forwards, because virtual stages 0 through $p-1$ are the first chunk on devices 0 through $p-1$. But each of those forwards is now $1/v$ of a full one. The warm-up costs $v$ times less time:

\[\text{bubble} = \frac{1}{v}\cdot\frac{p-1}{m}\]

The fourth step of the figure zooms into the warm-up of both schedules, which is where the entire difference lives, and measures the bubble off each.

The price is $v$ times as many point-to-point messages, over the slowest links in the job. At 16.8 MB a message that is a trade almost anyone would take. Meta used interleaving for Llama 3 and quoted the bubble in the paper as $(PP-1)/(V \cdot M)$.

What does splitting the backward pass buy?

A backward pass is two computations glued together. One computes the gradient with respect to the layer’s input, which the previous stage is blocked on. The other computes the gradient with respect to the layer’s weights, which nobody is blocked on.

The Zero Bubble paper’s whole idea is to stop treating them as one unit. Call them $B$ and $W$. Now $B$ is half as long, so the critical path through the pipeline is shorter, and $W$ can be dropped into any gap after its own $B$.

Their ZB-H1 schedule keeps 1F1B’s peak memory exactly and takes the bubble from $(p-1)(F+B+W)$ to $(p-1)(F+B-W)$, which is a third of the size when the three passes take equal time. The fifth step of the figure shows the ochre $W$ blocks filling what were holes, and measures 18.8% falling to 6.3%.

ZB-H2 goes further, adding forwards to the warm-up and reordering the $W$s at the tail so the shape becomes a parallelogram rather than a trapezoid. That reaches $(p-1)(F+B-2W)$, which is zero when the three take the same time, and it costs about twice the activations in flight. Getting the last of the bubble also required removing the synchronisation in the optimizer step, which the paper handles with a post-hoc validation rather than an up-front all-reduce.

DeepSeek-V3 went further again with DualPipe: two pipelines running in opposite directions at once, with a bubble of $(p/2 - 1)(F\&B + B - 3W)$, paid for with two copies of the parameters. They built it because their expert-parallel all-to-all needed something to hide behind, which is Part 5’s problem.

How much memory is actually in flight?

This is the part people get wrong, and it is a good interview question because the naive answer is so tempting.

Deepening the pipeline divides the weights by $p$. It does not divide the activations, because with a bubble-minimising schedule the first stage has to keep $p$ micro-batches in flight to stay busy, and each of those carries $L/p$ layers of activations.

\[p \times \frac{L}{p} = L \text{ layers' worth, for every } p\]

Korthikanti and colleagues state it exactly that way: the total activation memory on the first stage is $\frac{s\,b\,h\,L}{t}\left(34 + \frac{5as}{h}\right)$, with no $p$ in it at all. On our shape that is 22.8 GB regardless of whether the pipeline is four stages deep or sixteen.

Interleaving makes it slightly worse, by a factor of $1 + \frac{p-1}{pm}$, because each device holds pieces of two places in the network. Nobody notices that term; everyone notices the $L$.

How do you choose the number of micro-batches?

You do not choose it directly. It falls out of a batch identity:

\[m = \frac{\text{global batch}}{\text{micro-batch} \times d}\]

Larger $m$ shrinks the bubble as $1/m$. But at a fixed global batch, every GPU added to the data-parallel dimension takes micro-batches away from the pipeline. That is why the pipeline degree and the data-parallel degree cannot be chosen independently, and why frontier runs hold tokens per batch constant and accept a falling MFU rather than letting the bubble grow.

The last step of the figure plots the bubble against $m$ for one, two and four-way interleaving. At four stages you need about thirty micro-batches for a 10% bubble with plain 1F1B, or eight with four-way interleaving.

And a practical constraint that a lot of implementations impose and Meta had to remove: many schedules require the micro-batch count to be divisible by the number of stages. Their fix for Llama 3 was to make the number of consecutive micro-batches per stage a free parameter, so they could run fewer micro-batches than stages when the batch was tight, or more when they wanted to hide point-to-point latency.

What else goes wrong in a real pipeline?

Two imbalances that the clean picture does not show, and Meta documented both.

The first stage carries the embedding table on top of its layers, and carries the most warm-up micro-batches, so it uses the most memory. The last stage computes the output projection and the loss, so it is the slowest, and in a pipeline the slowest stage sets the pace for everyone. Their fix was to take one transformer layer away from each end: the first chunk holds only the embedding, the last only the projection and the loss.

Beyond that, the list is unglamorous. Asynchronous point-to-point so a stage does not block on a send. Proactive deallocation of stage input and output tensors that nothing will read again. And a document mask that makes different micro-batches cost different amounts, which is a straggler problem with a pipeline shape.

With all of that, Meta reported pre-training Llama 3 on 8192-token sequences with no activation checkpointing at all, which is the outcome the whole exercise is for.

They will ask Your pipeline is 16 stages deep and the profiler shows long idle bands at the ends of every step. What do you do?
The 30-second version First put a number on it: fifteen over the micro-batch count is the expected bubble, so at eight micro-batches it is 188% of the ideal time and the profile is behaving exactly as the formula says. Then check whether the micro-batch count is small because the global batch is small or because the data-parallel degree is large, since those have different fixes. If micro-batches are available, raise them, since 1F1B caps in-flight activations at the pipeline depth so it costs nothing in memory. If they are not, interleave: two or four chunks per device divides the bubble by that factor for the price of proportionally more point-to-point messages. If the schedule is already interleaved, split the backward pass and use a zero-bubble schedule, which takes another factor of three with no memory cost. And check the two imbalances before any of that, because a slow last stage with the loss on it, or a first stage carrying the embeddings and every warm-up micro-batch, looks like a bubble in a summary and is not one.

Rapid fire: can you do these from memory?

  1. Derive the bubble fraction and say what it is at eight stages and sixteen micro-batches.
  2. Say what 1F1B changes relative to GPipe, and what it does not.
  3. Give the interleaved bubble formula and the cost that pays for it.
  4. Explain what B and W are and why splitting them shortens the critical path.
  5. State the peak activation memory of the first stage in layers, and why p does not appear in it.
  6. Write the identity that determines the micro-batch count, and say what competes for it.
  7. Name the two structural imbalances in a real pipeline and a fix for each.

Part 5 takes the two axes people reach for last, and the two that interviews probe hardest: splitting the sequence, and splitting the experts.