Splitting a model across stages
Start with the shape of the cut. A transformer is a stack of L identical layers. Pipeline parallelism partitions that stack into p contiguous groups called stages, one per device, each holding roughly L / p layers. Device i runs a forward pass on its layers, sends the activations to device i+1, and on the way back receives gradients from i+1 and sends its own to i-1.
The appeal is that per-device parameter memory drops by a factor of p: a 70B model that needs ~140 GB in fp16 for weights alone splits into ~20 GB slices across 8 stages. Only the activations at stage boundaries cross the wire, and that volume — micro_batch × seq_len × d_model per hand-off — is tiny next to the all-reduces tensor parallelism demands. That cheap, point-to-point communication is why PP scales even across nodes where tensor parallelism would choke.
The problem: a strict data dependency
The catch is that the stages are not independent — they are a chain. Stage i+1 cannot begin until stage i has produced its output, and the backward pass cannot begin on any stage until the forward pass has reached the very end and the loss is computed. If you feed the whole batch through as one unit, the timeline is brutal: stage 0 works while stages 1 through p-1 sit idle, then stage 1 works while the rest wait, and so on.
With naive whole-batch execution, only one of p devices is ever active. Utilization is 1/p — an 8-way pipeline running at 12.5 percent efficiency, worse than not splitting at all. Recovering throughput requires breaking the batch into smaller pieces so different stages work on different pieces at once. Those pieces are micro-batches, the lever the rest of the math turns on.
Micro-batches fill the pipe
Split the global batch into m micro-batches. Now stage 0 processes micro-batch 1 and immediately hands it to stage 1; while stage 1 works on micro-batch 1, stage 0 starts micro-batch 2. After a short fill period every stage is working on a different micro-batch simultaneously — the assembly line is full and all p devices run in parallel.
It is exactly CPU instruction pipelining: the first instruction takes several cycles to traverse, but once the pipe is full a result retires every cycle. Here the ‘instruction’ is a micro-batch and the ‘stages’ are devices. The steady state is efficient; the inefficiency lives entirely at the two ends — the fill, while the pipe fills toward full occupancy, and the drain, while the last micro-batches trickle out and early stages run dry. Those two triangles of idle time are the bubble, and their size relative to useful work is what we now compute.
Deriving the bubble fraction
Measure time in units of one stage processing one micro-batch (take forward and backward as one combined unit for now). The last stage cannot start its first micro-batch until that micro-batch has traversed the previous p-1 stages — that is the fill cost, p-1 units. Symmetrically, after the last stage finishes, the pipeline drains for another p-1 units as work clears the remaining stages.
useful work per device = m
fill + drain idle (bubble) = p - 1
total wall-clock time = m + (p - 1)
bubble fraction = (p - 1) / (m + p - 1)This single ratio governs pipeline efficiency; utilization is its complement, m / (m + p - 1). The structure is intuitive: the bubble grows with the stage count p (a longer pipe takes longer to fill and drain) and shrinks as micro-batches m grow (more work to amortize the fixed fill/drain cost against). Everything else in PP scheduling makes one of those two terms more favorable.
A worked example
Take an 8-stage pipeline, p = 8, and run it with m = 8 micro-batches — a common first guess of one micro-batch per stage. The bubble fraction is (8-1)/(8+8-1) = 7/15 ≈ 0.47. Nearly half the pipeline’s wall-clock time is idle. That is a disaster hiding behind a plausible-looking configuration.
Now raise m to 32: 7/39 ≈ 0.18. At m = 64 it is 7/71 ≈ 0.10. To push under 5 percent you need m ≈ 133. The rule of thumb: you want m to be several times p — often m ≥ 4p to 8p — before the bubble stops dominating. Below that, adding pipeline stages can actively slow you down, because each new stage lengthens the fill and drain while the micro-batch count stays put. The number of micro-batches, not the number of GPUs, decides whether PP is worth using at all.
GPipe: all forward, then all backward
The original GPipe schedule is the simplest to reason about: push all m micro-batches forward through the pipeline, then run all m backward passes. Its bubble fraction is exactly the (p-1)/(m+p-1) derived above, and it is easy to implement because forward and backward phases are cleanly separated.
The cost is memory. Because backward for micro-batch 1 does not start until every forward pass is done, GPipe must keep the activations of all m in-flight micro-batches stashed for the backward pass. Peak activation memory therefore scales with m — and m is exactly the quantity you wanted to make large to shrink the bubble. This is the central tension of pipeline parallelism: the bubble pulls m up, and memory pulls it down. GPipe blunts the memory side with activation re-computation (checkpointing), but the fundamental O(m) growth remains.
1F1B: one forward, one backward
The 1F1B schedule (one-forward-one-backward, from PipeDream and adopted by Megatron-LM) attacks the memory problem without changing the bubble fraction. After the fill phase, each stage alternates: do one forward micro-batch, then immediately do one backward micro-batch, in steady lockstep. Because a backward pass runs as soon as possible, the activations it needs are freed early instead of piling up.
The result is that the number of in-flight micro-batches whose activations must be retained is capped at roughly the pipeline depth p, not the micro-batch count m. Peak activation memory becomes O(p) — independent of m. This is the quiet workhorse result of modern PP: since memory no longer grows with m, you are free to crank m as high as you like. 1F1B delivers the same (p-1)/(m+p-1) efficiency as GPipe while removing the memory ceiling that stopped you from reaching it — which is why 1F1B, not GPipe, is the default in production frameworks.