The underflow problem in one sentence
FP16 (half precision) is attractive for training: it halves memory and, on tensor-core hardware, roughly doubles throughput versus FP32. But it buys that with a narrow dynamic range. The catch is not the loss or the activations — those sit comfortably in FP16’s range. It is the gradients. During backpropagation, gradients of the loss with respect to early-layer weights are products of many small numbers, and they routinely land in the range 1e-8 to 1e-10.
FP16 cannot represent numbers that small. Anything below roughly 2^-24 ≈ 6e-8 rounds to exactly zero. So a gradient that should nudge a weight by a hair instead reads as no signal at all — the weight is frozen. Multiply this across the many parameters whose gradients live in that tiny range and a large fraction of the network stops learning. The model stalls or diverges, and no amount of tuning the learning rate fixes it, because the information was destroyed the moment the gradient was rounded to zero. Loss scaling exists to rescue those numbers before rounding happens.
The FP16 number line and the underflow cliff
A FP16 value is 1 sign bit, 5 exponent bits, and 10 mantissa bits. Work the range out from the exponent field:
FP16 = [sign:1][exponent:5][mantissa:10], bias = 15
largest normal = 2^15 * (2 - 2^-10) ≈ 65504
smallest normal = 2^-14 ≈ 6.10e-5
smallest subnormal = 2^-14 * 2^-10 = 2^-24 ≈ 5.96e-8
underflow cliff = anything < ~2^-25 rounds to 0Below the smallest normal (2^-14), FP16 keeps going with subnormals — values with a leading zero that trade mantissa precision for a little extra reach, down to 2^-24. Past that is the cliff: there is no representable positive number between 2^-25-ish and zero, so everything there collapses to zero. A gradient of 3e-9 is roughly 2^-28 — well over the edge, and gone. The whole game of loss scaling is to shift the gradient distribution to the right, off the cliff, using the headroom that sits unused up near 65504.
The trick: scale the loss, and the chain rule does the rest
The elegance is that you do not touch the gradients directly. You multiply the scalar loss by a constant S before calling backward(). Because differentiation is linear, that single factor propagates uniformly to every gradient in the graph:
L’ = S · L
∂L’/∂w = ∂(S · L)/∂w = S · (∂L/∂w) for every parameter wEvery gradient is now S times larger than it would have been — the entire distribution slides up by the same factor. Choose S = 2^15 = 32768 and a gradient that was 2^-27 (underflowed to zero) becomes 2^-27 × 2^15 = 2^-12 ≈ 2.4e-4, comfortably inside FP16’s normal range and represented with full precision. Crucially, using a power of two makes the scaling exact in floating point — it only changes the exponent, never the mantissa, so no rounding error is introduced by the scaling itself. The gradients are now safe to store in FP16 through the backward pass.
Unscale before you step — and before you clip
Scaled gradients are correct for surviving the backward pass, but they are S times too big for the actual weight update. If the optimizer stepped on them directly, the effective learning rate would be S times too large and training would explode. So you unscale: divide the gradients by S before the optimizer step, typically as they are copied into the FP32 master weights the optimizer maintains.
The subtle part is gradient clipping. Clipping thresholds a gradient by its true norm (say, clip to norm 1.0). If you clip the scaled gradients, the norm is S times too large and the clip fires far too aggressively — every step gets crushed to the threshold. So the canonical order is fixed: backward(S·L) → unscale grads → clip on the unscaled grads → optimizer.step(). Anything that reads the gradient magnitude — clipping, gradient-norm logging, weight decay computed on gradients — must see the true, unscaled values. Scaling is a temporary disguise the gradients wear only for the trip through FP16 storage.
Static loss scaling: pick one number
The simplest scheme is static (fixed) loss scaling: choose a single constant S — commonly 128, 1024, or 2^15 — and use it for the whole run. It works when it works, and it is trivial to implement. The difficulty is picking the value, because it is a tightrope between two failure modes.
Set S too low and small gradients still underflow to zero — you have not lifted the distribution far enough off the cliff. Set it too high and the large gradients at the top of the distribution overflow: grad × S > 65504 becomes inf, which then poisons the whole update with inf/nan. The safe window depends on the model and even drifts during training as gradient magnitudes shrink toward convergence — a scale that was fine at step 0 may start overflowing later, or become too conservative. Static scaling therefore demands manual tuning and offers no safety net, which is exactly why the dynamic scheme was invented.
Dynamic loss scaling: let S find itself
Dynamic loss scaling removes the guesswork by treating S as a value the training loop discovers automatically. The principle: push S as high as possible to protect the smallest gradients, and back off only when the largest gradients overflow. Two rules run every step:
Backoff on overflow. After the backward pass, inspect the scaled gradients for inf or nan. If any appear, S was too high this step: skip the optimizer step entirely (the gradients are corrupt — do not update the weights) and multiply S by a backoff factor, usually 0.5. Growth on a clean streak. If some number of consecutive steps (often 2000) pass with no overflow, gently raise S by a growth factor, usually 2, and reset the counter. The result is a self-tuning value that hovers just below the overflow threshold — the largest scale that is currently safe — and re-adapts for free as gradient magnitudes change over training. This is what frameworks ship by default.