Why architecture matters here
Start with the arithmetic, because it is what makes the problem urgent rather than academic. A pipeline with p stages and m microbatches spends p-1 microbatch-times filling and p-1 draining. The bubble fraction is (p-1)/(m+p-1), commonly quoted as (p-1)/m when m is large. At p=8, m=16: 44%. At p=16, m=32: the same 44% — the ratio is what matters, not the absolute numbers. To push the bubble under 10% you need m ≈ 10p: 80 microbatches for 8 stages. Now count memory: GPipe stores activations for every in-flight microbatch, so 80 microbatches of activations must be resident simultaneously. That is exactly the memory you did not have. The bubble and the memory ceiling are the same constraint viewed from two directions, and every schedule in this space is a different trade between them.
This is why the naive advice — 'just use more microbatches' — is not merely suboptimal but often impossible. And it gets worse at scale, because the global batch size is not a free parameter either: it is fixed by optimization considerations, since batches beyond a critical size stop improving convergence per token. So m is bounded above by memory and the batch budget is bounded above by convergence, and both squeeze from the same side. Meanwhile p is bounded below by the model's size — you need enough stages to fit the parameters at all. You do not get to choose your way out of the bubble.
The insight that unlocks the problem is that the dependency graph is finer-grained than the schedule assumes. Autograd hands you a backward pass as one call, and every pipeline schedule before zero-bubble inherited that granularity without examining it. But the chain rule for a linear layer y = xW gives two separate products: dL/dx = dL/dy · Wᵀ and dL/dW = xᵀ · dL/dy. They read the same incoming gradient and are otherwise completely independent. The first is what the upstream stage is blocked on. The second feeds only the optimizer, at the end of the iteration, and could as well be computed hours later. Treating them as one unit imposes a false dependency: it forces work off the critical path to run on the critical path.
Once you see the split, the schedule almost designs itself. The bubble is idle time on the critical path; W is work with no critical-path deadline; therefore put W in the bubble. The cost is that W's inputs — the saved activation x and the incoming gradient dL/dy — must stay resident until W runs, so deferring W defers a memory release. That is the real trade, and a much better one than the alternative: it buys pipeline efficiency with a bounded, controllable amount of memory rather than with the unbounded microbatch count GPipe demanded. It also explains why a true zero bubble needs a memory budget above 1F1B's.
The architecture: every piece explained
F, B, and W (top row). Three primitives, and the whole architecture is a claim about their dependencies. F is forward: consume the previous stage's activation, produce this stage's, save what backward will need. It is strictly ordered — stage i's F for microbatch j needs stage i-1's F for microbatch j. B computes dL/dx and is strictly ordered in the other direction: stage i's B for microbatch j unblocks stage i-1's B for microbatch j. W computes dL/dW and has exactly one ordering constraint: it must happen after the B that produced its incoming gradient, and before the optimizer step. Between those two points it floats freely. That freedom is the entire resource this architecture spends.
The schedule lineage (middle row). GPipe runs all F's then all B's: simple, and it holds every microbatch's activations simultaneously — bubble (p-1)/m, memory O(m). 1F1B interleaves so each stage alternates one forward with one backward once the pipe is full; the bubble is identical, but activation memory drops to O(p) because a microbatch's activations are freed as soon as its backward passes through. This is the crucial and often-missed point: 1F1B is a pure memory win, not a bubble win, and it is what made deep pipelines feasible at all. Interleaved 1F1B gives each device v non-contiguous chunks of layers, cutting the bubble to (p-1)/(m·v) at the price of v× the point-to-point communication — a real trade, and one that stops paying when your interconnect is the bottleneck. Zero-bubble then splits B into B and W and uses W as filler.
Warmup, steady state, cooldown (lower-left). Every pipeline schedule has three phases, and knowing which one your bubble lives in tells you which fix applies. Warmup: stages start staggered, so stage i idles for i microbatch-times while the first activations propagate. Steady state: every stage has work every slot — a well-formed 1F1B steady state is already bubble-free, which surprises people. Cooldown: the drain, symmetric to warmup. So the bubble is entirely in the tails, and zero-bubble's real job is filling the tails with W work: during warmup a stage has no B to do yet but has accumulated W from earlier microbatches; during cooldown it has no F left but plenty of deferred W. The handcrafted ZB-H1 and ZB-H2 schedules are precisely the answer to 'how much W do I defer, and to where', with ZB-H1 matching 1F1B's memory and roughly halving the bubble, and ZB-H2 reaching a true zero bubble at a memory cost above 1F1B.
The optimizer sync bubble (lower-right). After all this, one bubble survives, and it is not about pipelining at all. The optimizer step traditionally begins with a synchronization: an all-reduce of gradient norms for clipping, plus a check for NaN or inf in mixed-precision training to decide whether to skip the step. Every stage must wait for a global answer before updating, and that barrier is a bubble at the end of every single iteration. Zero-bubble's post-validation strategy removes it by inverting the order: optimistically perform the step, then validate afterward, and roll back on the rare iteration where validation fails. Since inf/NaN steps are rare, the expected cost of occasional rollback is far below the guaranteed cost of synchronizing every iteration — a classic optimistic-concurrency trade, applied to a place nobody was looking for one.
End-to-end flow
Warmup. Iteration begins. Stage 0 runs F for microbatch 1 and ships the activation to stage 1, then F for microbatch 2, and so on — it has work immediately. Stage 7, at the far end, has nothing at all: no activation has reached it. Under 1F1B, stage 7 simply idles for seven microbatch-times. This is the warmup bubble, and it is pure waste — 7 GPUs idle while the pipeline fills. Under zero-bubble the deep stages still idle during the very first iteration (there is genuinely no work in existence yet), but from the second iteration onward there is deferred W from the previous iteration's tail available to fill the slot, which is why zero-bubble's benefit is measured in steady-state iterations rather than in a single-step microbenchmark.
Steady state. The pipe is full. Stage 4 alternates: F for microbatch 12, B for microbatch 8, F for 13, B for 9. Each B produces dL/dx, which is immediately sent upstream to stage 3 — that is the critical path and it must not wait. Each B also leaves behind a pending W: the saved activation and incoming gradient sit in memory with a note that dL/dW has not been computed yet. The queue of pending W grows. Notice that the steady state has no bubble to fill — every slot is busy — so the deferred W is not filling anything here; it is being banked for the cooldown, where the bubble actually lives. The scheduler's job in steady state is simply to defer, and to defer no more than the memory budget allows.
Cooldown. Stage 0 has finished all its forwards and all its B's. Under 1F1B it would now go idle for seven microbatch-times while the backward wave completes on downstream stages — the cooldown bubble, the mirror image of warmup. Under zero-bubble, stage 0 opens its pending-W queue and starts computing weight gradients: dL/dW for microbatch 1, then 2, then 3. This work was always going to happen; it has simply been moved from the critical path into a slot that was going to be idle. Stage 0's utilization goes from 0% to 100% for the duration, and every stage does the same as its own backward wave passes. The bubble has been filled with work rather than eliminated by a trick.
The step. All W's are drained, so every weight gradient exists. Traditionally the optimizer would now all-reduce gradient norms and check for inf/NaN before deciding to apply — one global barrier, one final bubble. With post-validation, each stage applies its update immediately and validation runs asynchronously behind it; if the check later reports an inf, the step is rolled back using the saved optimizer state. The rollback costs one wasted iteration, and it happens on well under 1% of steps in a healthy run, so the expected cost is a rounding error against a barrier paid every iteration. Net result: an 8-stage pipeline that spent 44% of its time idle under GPipe now runs near-continuously, with 1F1B's activation-memory ceiling plus a bounded allowance for deferred W.