Why the optimizer state is the biggest bucket

Four things compete for GPU memory during training: parameters, gradients, optimizer state, and activations. Activations can be traded away with recomputation; the other three scale purely with the parameter count Ψ. Count bytes per parameter for mixed-precision Adam. The GPU holds an fp16 weight (2 bytes) and an fp16 gradient (2 bytes). The optimizer, working in fp32 for stability, holds three fp32 tensors: a master copy of the weights (4 bytes), the first moment m (4 bytes), and the second moment v (4 bytes).

TensorPrecisionBytes/paramBucket
Weightfp162parameters
Gradientfp162gradients
Master weightfp324optimizer
Momentum mfp324optimizer
Variance vfp324optimizer

That is 16 bytes/param, of which 12 — the fp32 master, m, and v — are optimizer state. Optimizer state is 12/16 = 75% of the total. Evicting it is the single biggest memory win available.

Advertisement

What Adam is actually storing

The momentum m and variance v are exponential moving averages that give Adam its per-parameter adaptive step. For gradient g_t:

m_t = β1 · m_(t-1) + (1 - β1) · g_t
v_t = β2 · v_(t-1) + (1 - β2) · g_t^2
m̂ = m_t / (1 - β1^t)      v̂ = v_t / (1 - β2^t)
θ_t = θ_(t-1) - α · m̂ / (√v̂ + ε)

Two things matter for offload. First, m and v are stateful: they must persist across steps, so they cannot simply be recomputed like activations can. Second, the update is element-wise — each parameter’s new value depends only on its own g, m, and v. There is no matrix multiply, no cross-parameter coupling. That makes the step cheap in FLOPs and embarrassingly parallel, which is exactly what lets a CPU keep up.

Advertisement

The offload idea in one sentence

Keep the forward and backward passes on the GPU; move the optimizer state — fp32 master weights, m, and v — into CPU DRAM, and run the parameter-update step on the CPU itself. The GPU computes fp16 gradients and ships them over PCIe to the host; the CPU folds them into m and v, applies the Adam update to the fp32 master weights, and ships the updated fp16 weights back for the next forward pass. The fp32 state never touches the GPU — the accelerator only ever holds fp16 weights and gradients, 4 bytes/param instead of 16. The 12 bytes of optimizer state live entirely on the host, so GPU memory for weights-and-state drops by a factor of four.

How ZeRO-Offload partitions compute and communication

Why split the work at the optimizer step and nowhere else? ZeRO-Offload models training as a data-flow graph and asks where to draw the CPU/GPU boundary to minimize two costs at once: PCIe traffic and CPU compute. The forward and backward passes cost O(Ψ · B) FLOPs — they scale with both the parameter count and the batch size B, so they are compute-heavy and belong on the GPU.

The optimizer step, by contrast, costs only O(Ψ) FLOPs — a handful of element-wise operations per parameter, independent of batch size. It is memory-bound, not compute-bound. Assigning that low-FLOP node to the CPU moves the entire 12-byte state off the GPU while adding the least possible compute to the host. Any other cut — offloading part of the backward pass, say — would either move far more FLOPs to the slow CPU or move far more bytes across PCIe. The optimizer step is the unique sweet spot.

The communication budget

Offload is only a win if the PCIe transfer does not swamp the savings. Per step, the traffic is fixed: fp16 gradients travel GPU→CPU and updated fp16 weights travel CPU→GPU. That is 2Ψ + 2Ψ = 4Ψ bytes crossing the bus, once per iteration.

The key property is that this 4Ψ is independent of batch size, while GPU compute grows as O(Ψ · B). So the ratio of transfer time to compute time falls as the batch grows — a large enough batch amortizes the PCIe cost until it hides under the backward pass. This is why optimizer offload pairs naturally with large batches and gradient accumulation: each expensive compute phase is stretched long enough that the constant communication tax rounds to zero. Run tiny batches and the same fixed transfer becomes a visible bottleneck.

A worked example: a 7B model

Take Ψ = 7×10^9 parameters. On-GPU without offload, weights-plus-state costs 16 · Ψ ≈ 112 GB — already past a single 80 GB accelerator before a single activation is stored.

With optimizer-state offload, the GPU keeps only the fp16 weights and fp16 gradients: 4 · Ψ ≈ 28 GB. The 12-byte optimizer state — 12 · Ψ ≈ 84 GB — moves to host DRAM, where 84 GB is unremarkable. Per step, PCIe carries 4 · Ψ ≈ 28 GB. On a 16 GB/s PCIe 3.0 link that is roughly 1.75 s of raw transfer; on a bidirectional PCIe 4.0 link, closer to 0.9 s — tolerable only because a large-batch forward/backward on a 7B model already takes seconds, so with overlap the transfer largely hides beneath it. The 84 GB you no longer need on the GPU is the payoff.

CPU-Adam: making the host step fast enough

The obvious worry is that a naive CPU optimizer step over billions of parameters is slow enough to stall the GPU. ZeRO-Offload answers this with CPU-Adam, a hand-tuned host implementation that keeps the step off the critical path.

It leans on three things. SIMD vectorization (AVX2/AVX-512) processes 8–16 fp32 lanes per instruction, matching the element-wise structure of the update. Loop tiling keeps each parameter’s m, v, master weight, and gradient resident in cache together, so the memory-bound step streams at DRAM bandwidth rather than thrashing. Multithreading fans the independent per-parameter work across every core. Because the update is O(Ψ) and perfectly parallel, a modern multi-core CPU sustains billions of parameter updates per second — fast enough that, overlapped with GPU compute, the optimizer step adds little wall-clock time.