Pipeline Schedules: GPipe, 1F1B and Zero Bubble
The problem GPipe leaves behind
The previous note derived the bubble fraction $(P-1)/(m+P-1)$ and showed that shrinking it needs many microbatches — $m\ge 9(P-1)$ for a 10% bubble. It also flagged the cost: under GPipe's all-forward-then-all-backward ordering, stage 0 holds activations for all $m$ microbatches simultaneously.
So the two goals are in direct conflict. Raise $m$ to kill the bubble and you raise memory by the same factor. At $P=8$ and $m=64$ that is 64 microbatches of activations on stage 0, which no accelerator will hold.
The resolution is that nothing requires all forwards to precede all backwards. A microbatch's backward pass can run as soon as its forward has reached the last stage and come back. Running it earlier frees its activations earlier.
1F1B: same bubble, bounded memory
The one-forward-one-backward schedule has three phases:
Warm-up. Stage $p$ runs $P-1-p$ forward passes to fill the pipe. Earlier stages run ahead; the last stage runs none.
Steady state. Each stage strictly alternates: one forward, one backward, one forward, one backward. Every backward immediately frees the activations of the microbatch it consumed, so memory stops growing.
Cool-down. The remaining backward passes drain out.
Count the microbatches in flight at stage $p$: it is $P-p$, so the worst case is stage 0 with $P$. That is the whole result:
This is why 1F1B, not GPipe, is the default in every serious framework. GPipe remains the clearer teaching example, and the two are identical in throughput.
Interleaved stages
1F1B fixed memory but left the bubble alone. To attack the bubble itself, give each device several non-contiguous chunks of the model instead of one contiguous block. With $v$ chunks per device — Megatron calls $v$ the virtual pipeline size — the pipeline has $vP$ logical stages running on $P$ physical devices, and the fill and drain overlap more finely:
| Configuration | $v=1$ | $v=2$ | $v=4$ |
|---|---|---|---|
| $P=4$, $m=8$ | 27.3% | 15.8% | 8.6% |
| $P=8$, $m=8$ | 46.7% | 30.4% | 17.9% |
| $P=8$, $m=16$ | 30.4% | 17.9% | 9.9% |
Zero bubble: split the backward pass
Every schedule so far treats the backward pass as one indivisible unit. It is not. From the FLOPs note, backward consists of two independent matmuls:
B — the input gradient
On the critical path. The previous stage cannot start its own backward until this arrives.
W — the weight gradient
Nothing downstream depends on it. It only has to be finished before the optimizer step at the end of the batch.
Handling the fill and drain differently gives a family of schedules: heuristic variants that reduce the bubble by roughly a third to a half, and a fully optimized variant that eliminates it almost entirely at the price of extra memory and a more complex optimizer step.
Choosing a schedule
| Schedule | Bubble | Peak activations | P2P traffic | Use when |
|---|---|---|---|---|
| GPipe | $\dfrac{P-1}{m+P-1}$ | $O(m)$ | baseline | teaching; almost never in production |
| 1F1B | $\dfrac{P-1}{m+P-1}$ | $O(P)$ | baseline | the sensible default |
| Interleaved 1F1B | $\dfrac{P-1}{vm+P-1}$ | $O(vP)$ | $v\times$ | bubble hurts and the interconnect has headroom |
| Zero bubble | near zero | $O(P)$ plus gradient buffers | baseline | large $P$, and you can afford the machinery |
One more lever sits outside the schedule entirely: activation recomputation. Storing only each stage's inputs and re-running the forward pass during backward makes GPipe's memory almost independent of $m$ too, at a 33% compute surcharge. It composes with any schedule, and combining 1F1B with selective recomputation is the usual production answer.
1F1B gives you GPipe's throughput with memory proportional to $P$ instead of $m$ — take it unconditionally. Interleaving trades communication for a smaller bubble. Zero-bubble schedules exploit the fact that the weight gradient has no downstream dependency and can therefore be deferred into the idle slots.
Check yourself
- Show that stage $p$ holds $P-p$ microbatches in flight under 1F1B. derivation
- Explain why 1F1B has the same bubble fraction as GPipe despite reordering. reasoning
- Compute the interleaved bubble for $P=16$, $m=32$, $v=4$. calculation
- Explain why $\bar W$ can be deferred but $\bar X$ cannot. derivation
- Compare 1F1B plus recomputation against interleaved 1F1B for a pipeline spanning a slow network. analysis