Three levers on one budget
Training memory splits into two piles. The static pile — weights, gradients, optimizer state — is fixed by the model and is what ZeRO/FSDP sharding attacks. The dynamic pile is activations: every intermediate tensor autograd must keep until the backward pass reaches it. The dynamic pile is the one that scales with your batch and your context length, and it is almost always what actually blows up.
Without changing the model, there are exactly three ways to shrink it. You can make less of it (smaller micro-batch), throw it away and rebuild it (activation recomputation), or spread it across devices you already own (sequence parallelism). These are not equivalent. Two of them cost you throughput; one is close to free. Knowing which is which, and how far the free one reaches before it runs out, is the whole practical skill — and it is a different skill from being able to derive the formula.
What the knob buys, stated once
The activation count for one transformer layer, with sequence length s, micro-batch b, hidden size h, and a heads, follows Korthikanti et al. (2022). Written for a single device, for tensor parallelism at degree t, and for TP+SP:
A_1 = sbh · (34 + 5as/h) one device
A_TP = sbh · (10 + 24/t + 5as/(h·t)) TP alone, degree t
A_TP+SP = sbh · (34 + 5as/h) / t TP + sequence parallel
saving = 10·sbh · (1 − 1/t) per layer, per deviceRead the last line as the knob’s entire range. SP does not scale with t the way you might hope: it removes the replicated 10·sbh floor and nothing else, and even at t = ∞ that is all it can ever remove. The interesting question is therefore not ‘how much does SP save’ in the abstract, but what fraction of your particular total is that floor — a number that swings from trivial to dominant depending on the rest of your configuration.
Free on the wire, not free in sync points
The reason SP is described as free is that a ring all-reduce is already implemented as a reduce-scatter followed by an all-gather. SP just uses the two halves separately — all-gather entering a matmul region, reduce-scatter leaving it — so both the bytes moved, 2(t−1)/t · V per device, and the 2(t−1) ring steps are unchanged. Bandwidth-neutral and, in the ring model, latency-neutral too.
What does change is the count of collective calls. A TP layer issues two collectives; a TP+SP layer issues four. Same total traffic, twice as many synchronization points. In practice that shows up as extra kernel-launch overhead and slightly more exposure to stragglers and network jitter, because every collective is a barrier where the fastest rank waits for the slowest. It is a small tax, typically a low single-digit percentage, and it is why SP is ‘nearly free’ rather than free — but it is real, and it grows on noisy or oversubscribed fabrics.
The knob you cannot turn on its own
Here is the constraint that surprises people configuring a run: in Megatron-style SP there is no independent SP degree. The whole mechanism is a reshard between two views of the same t ranks — sequence-split in the norm band, hidden-split in the matmul band. SP degree is TP degree, always. In the config it appears as a boolean, not a number.
The consequence is that you cannot buy more activation savings by raising SP alone. Wanting a bigger divisor means raising t, which drags along all of tensor parallelism’s costs: more all-reduce traffic per layer, and the hard requirement that the TP group stay inside one NVLink domain. Two smaller constraints follow. s must be divisible by t, so odd context lengths need padding that you must then mask out of the loss. And after the reshard each rank feeds b · s/t tokens into its GEMMs — drive that too low and the matmuls go skinny and lose efficiency.
Context parallelism is the axis that is independent
SP is bounded because it only touches the token-independent band: LayerNorm normalizes each token across the hidden dimension and residual dropout is elementwise, so neither mixes positions and both shard for free. Attention is the opposite — every query must see every key — so SP leaves the attention core alone.
Context parallelism (ring attention) is the separate knob that does shard attention, by rotating K/V blocks around a ring so no device ever holds the full sequence; the mechanism is covered separately. The configuration point is that CP is a real mesh dimension with its own degree: N = DP × CP × TP × PP. So the two sit in a clean relationship. SP is free but bounded — it costs nothing and caps out at 10·sbh. CP is unbounded but not free — it scales to any context but adds a per-layer K/V ring and a causal-mask load-balance problem. Reach for SP first; reach for CP when the attention core, not the norm band, is what overflows.
SP versus simply using a smaller micro-batch
Every activation term is linear in b, so halving the micro-batch halves all of them — including the quadratic attention term that SP cannot reach. On pure memory-per-dollar, shrinking b is the stronger lever. So why is it the second choice?
Because it is the one that costs throughput. Holding the global batch fixed (it is set by convergence, not by your memory), a smaller micro-batch means proportionally more gradient-accumulation steps. Each step re-pays the fixed per-step overheads, and more importantly the GEMMs get narrower: arithmetic intensity falls, the matmuls drift from compute-bound toward memory-bound, and utilization sags well before you reach b = 1. SP has no equivalent penalty — it moves bytes you were already moving and leaves every GEMM the same shape. The rule that falls out: exhaust the free lever before you pay for the expensive one. Turn SP on, then shrink b only for the memory SP could not reach.