Rounding is not free in low precision

A floating-point format represents only a discrete grid of values. The spacing between adjacent representable numbers near a value x is one ulp (unit in the last place), and it grows with magnitude: near 1.0 the grid is fine, near 256.0 it is coarse. Whenever a computation produces a number not exactly on the grid — almost always — it must be rounded to a neighbour before it can be stored.

In FP32 the ulp is so small that rounding error is negligible for most training. In BF16, FP8, and FP4 it is not: the mantissa is short, so the grid is coarse and each rounding discards a meaningful number of bits. The question of how you round — not just to how many bits — becomes a first-order design choice. Stochastic rounding is one answer, and it targets a specific, damaging failure of the obvious default.

Advertisement

Round-to-nearest and the swamping problem

The default everywhere is round-to-nearest (RTN): given x between grid points x_lo and x_hi = x_lo + ulp, pick whichever is closer. RTN minimises the error of a single rounding, and for one-shot quantisation that is exactly what you want. The trouble is what it does to a running sum.

Consider adding a small update g to a large weight w. If |g| < ulp/2, then w + g is still closer to w than to any other grid point, so RTN rounds it straight back to w and the update vanishes. This is swamping (or stagnation): a small quantity added to a large one is annihilated by rounding. Repeat it every step — exactly what SGD does — and the weight never moves, however many nonzero gradients arrive. The systematic bias does not average out; it accumulates into a training that stalls. Small learning signals, precisely the ones fine-tuning depends on, are the first casualties.

Advertisement

Stochastic rounding, defined

Stochastic rounding (SR) replaces the deterministic ‘pick the closer one’ with a coin flip whose bias depends on position. For x lying between x_lo and x_hi = x_lo + ulp, round up to x_hi with a probability equal to how far x has travelled from x_lo toward x_hi, and otherwise round down to x_lo:

x_lo = floor(x / ulp) * ulp        (grid point just below x)
x_hi = x_lo + ulp                 (grid point just above x)

                x_hi ,  with probability  p = (x - x_lo) / ulp
SR(x)  =  {
                x_lo ,  with probability  1 - p

The closer x sits to x_hi, the more likely SR rounds up; the closer to x_lo, the more likely it rounds down. If x is already on the grid, p = 0 and SR leaves it alone. Crucially, SR needs the extra low-order bits of x — the part being discarded — to compute p, so it applies when a higher-precision intermediate is cast down.

The probability formula

The rounding probability is exactly the fractional part of x measured in units of ulp:

p = (x - x_lo) / ulp = frac(x / ulp)   where  frac(y) = y - floor(y)

This holds because x_lo = floor(x/ulp) · ulp, so dividing x - x_lo by ulp leaves precisely x/ulp - floor(x/ulp). The probability lives in [0, 1): at the lower grid point p = 0 (never round up), and as x approaches the upper point p → 1 (almost always round up). At the midpoint p = 0.5 — a fair coin.

In practice you do not compute a floating-point division. Hardware draws a uniform random value from the bits being truncated and compares: if u ~ Uniform[0,1) satisfies u < p, round up, else round down — one random draw and a comparison per element, cheap enough to sit inside a tensor cast.

Why it is unbiased

The whole point of the position-dependent coin is that the expected rounded value equals the input exactly. With a two-point distribution, the expectation is the probability-weighted average of the two outcomes:

E[SR(x)] = p · x_hi + (1 - p) · x_lo
         = x_lo + p · (x_hi - x_lo)
         = x_lo + p · ulp
         = x_lo + (x - x_lo)          (since p · ulp = x - x_lo)
         = x

So E[SR(x)] = x: stochastic rounding is an unbiased estimator of the true value. Round-to-nearest is not — whenever it snaps w + g back to w, its expected output is w, a systematic error of -g that never cancels. Better still, unbiasedness composes across steps: by the tower property of expectation, a chain of unbiased rounds has an expected trajectory equal to the exact, infinite-precision one. The individual weights are noisy; their expected path is right.

A worked example: the weight RTN freezes

Take a weight stored on a grid with ulp = 1 at this magnitude, sitting at w = 256, and a per-step update of g = 0.3. Each step computes w + g = 256.3 in higher precision, then must round back to the grid.

x = 256.3   x_lo = 256   x_hi = 257   ulp = 1
p = frac(256.3 / 1) = 0.3

RTN:  0.3 < 0.5  →  round to 256   (update lost, every single step)
SR:   round to 257 with prob 0.3, else 256

Under RTN the weight is stuck at 256 forever — the gradient might as well be zero. Under SR, each step rounds up to 257 with probability 0.3. After 10 steps the exact answer is 256 + 10 × 0.3 = 259, and in expectation SR drifts the weight to 259 too. The tiny update that RTN annihilates is recovered on average, at the cost of a jittery path.

Why unbiasedness preserves gradient information

Stochastic gradient descent is already a noisy process — every mini-batch gradient is a random estimate of the true gradient, and training works precisely because those estimates are unbiased and average out over many steps. SR slots into this picture perfectly: it adds one more source of zero-mean noise while keeping the estimate unbiased, so the optimiser’s existing averaging absorbs it.

RTN, by contrast, injects a bias, and bias is the one thing SGD cannot average away. A consistent -g error every step is not noise that cancels; it is a systematic force pulling the update toward zero, and it wins whenever the update is sub-ulp. Small gradients — late-training fine adjustments, the tail of a learning-rate decay, subtle features — are exactly where updates fall below ulp/2. SR keeps those alive because their information lives in the probability of rounding up, not in any single rounded value.