The memory wall that makes offload necessary
Training a transformer needs far more memory than the parameters alone. Take a model with P parameters trained in mixed precision with Adam. The GPU must hold, per parameter: an fp16 weight (2 bytes), an fp16 gradient (2 bytes), and Adam’s bookkeeping in fp32 — a master copy of the weight (4 bytes), a first-moment estimate (4 bytes), and a second-moment estimate (4 bytes). That is 2 + 2 + 4 + 4 + 4 = 16 bytes per parameter before a single activation is stored.
So a 7 billion-parameter model needs roughly 7e9 × 16 = 112 GB just for weights, gradients, and optimizer state — impossible on a 24 GB card, and tight even on an 80 GB one once activations are added. The observation behind offload is that most of those 16 bytes are idle most of the time: the optimizer state is touched only during the update, and the gradient exists only after backward. Idle tensors do not need to sit in scarce GPU memory.
Splitting the budget: what to keep, what to offload
Group the 16 bytes by how often the GPU touches them. The fp16 weights (2 bytes) are read on every forward and backward pass — they are hot and belong on the GPU. Activations (not in the 16) are also hot. The other 12 bytes — the fp32 master weight and the two Adam moments — are touched only during the optimizer step. That is the natural offload target.
Pushing those 12 bytes per parameter to CPU RAM cuts the persistent GPU footprint from 16 to 4 bytes per parameter — a 4× reduction in the weight/grad/state budget. For the 7B model that is the difference between 112 GB and about 28 GB. This is the split popularized by ZeRO-Offload: fp16 parameters and gradients live on the GPU, the fp32 optimizer state lives on the CPU, and the Adam update itself runs on the CPU, so those 12 bytes never cross to the GPU at all — only the updated fp16 weights are copied back.
Where the update runs decides what crosses the bus
A subtle but decisive choice: do you offload the optimizer data and run the update on the GPU, or offload the update computation too? Keep the Adam step on the GPU and you must stream all 12 bytes of state up, update, and stream it back — roughly 24 bytes of traffic per parameter every step. Let the CPU run Adam and the GPU sends only the fp16 gradient down (2 bytes) and receives the updated fp16 weight back (2 bytes) — about 4 bytes per parameter.
That is a 6× difference in bus traffic, which is why practical offload schemes push the computation to the CPU. The trade is that the CPU is far slower at arithmetic, so the Adam step — memory-bound and embarrassingly parallel — must be well vectorized (AVX, multiple threads) or it becomes the new bottleneck. The math still favors it: an Adam update is a handful of FLOPs per parameter, trivial next to the matrix multiplies of forward and backward.
The PCIe bottleneck, in numbers
Everything offloaded must cross the PCIe bus, and PCIe is slow relative to on-GPU memory. Round numbers for a single x16 link:
PCIe 3.0 x16 ~16 GB/s
PCIe 4.0 x16 ~32 GB/s
PCIe 5.0 x16 ~64 GB/s
GPU HBM (A100) ~1500-2000 GB/s (on-device, for contrast)The GPU’s own memory is one to two orders of magnitude faster than the pipe to the CPU. That gap is why offload is delicate: you trade abundant-but-slow capacity for scarce-but-fast bandwidth. PCIe 4.0 moves about 32 bytes per nanosecond, so one gigabyte takes roughly 1e9 / 32e9 ≈ 31 ms. Multiply by the tens of gigabytes a large model moves per step and the transfer time becomes comparable to — or larger than — the compute time. Whether that kills throughput depends entirely on overlap.
The swap math: transfer time vs compute time
Reduce a training step to two competing durations. Let B be the bytes that must cross PCIe per step and BW the bus bandwidth; then transfer time is t_move = B / BW. Let the GPU’s useful work be F FLOPs at throughput R; then compute time is t_compute = F / R. If the two can proceed at the same time, the step takes max(t_compute, t_move); if they cannot overlap, it takes the sum.
Offload is ‘free’ precisely when t_move ≤ t_compute — the data you need next arrives before the GPU finishes the work it is already doing, so the GPU never stalls. Offload is fatal when t_move >> t_compute: the expensive GPU sits idle waiting for the pipe. The entire engineering effort of a good offload runtime goes into making the first inequality hold — shrinking B, and overlapping what remains.
Overlap: hiding the pipe behind the compute
A transformer is a stack of layers, and both forward and backward proceed layer by layer. That structure is what makes overlap possible. During the backward pass, the moment layer L produces its gradient, that gradient can start streaming to the CPU while the GPU is already computing the gradient of layer L-1. Likewise, the updated weights for a layer can be prefetched back onto the GPU before the next forward pass reaches that layer.
This is classic double buffering: a copy engine (the GPU’s DMA hardware, independent of the compute cores) moves tensor i across PCIe on one CUDA stream while the cores work on tensor i+1 on another. If each layer’s compute takes at least as long as moving its bytes, the transfers vanish into the shadow of the math and the step runs at nearly full GPU speed. Serialize the copy and the compute instead, and the same bytes cost a stall on every layer.
A worked example
Take a 7B-parameter model, PCIe 4.0 at 32 GB/s, CPU-side Adam. Per step the GPU sends fp16 gradients down and receives updated fp16 weights back: 2 + 2 = 4 bytes per parameter, so B = 7e9 × 4 = 28 GB. Transfer time is t_move = 28 / 32 ≈ 0.88 s per step (both directions counted).
Now the compute. A training step costs roughly 6 × P FLOPs per token (forward plus backward). With a batch of, say, 64k tokens: F = 6 × 7e9 × 64e3 ≈ 2.7e18 FLOPs. On a GPU sustaining 150 TFLOP/s of useful fp16, t_compute = 2.7e18 / 1.5e14 ≈ 18 s. Here t_move (0.88 s) << t_compute (18 s), so with overlap the offload traffic is completely hidden and costs almost nothing. Shrink the batch to 512 tokens and compute drops to about 0.14 s while transfer stays 0.88 s — now the bus dominates and throughput collapses. Same model, same hardware; the batch size decides whether offload is brilliant or ruinous.
The lever, then, is compute-per-byte-moved — a form of arithmetic intensity. Larger batches, longer sequences, and gradient accumulation all raise t_compute while leaving the per-step transfer B roughly fixed, so the counterintuitive rule is that offload gets more efficient the more work you push through each step. Accumulating gradients over many micro-batches before the optimizer update is the classic way to buy that headroom; small, latency-oriented steps are where offload hurts most.