The one-line idea: FP32 with the tail chopped off
A float32 lays out its 32 bits as 1 sign, 8 exponent, 23 mantissa. bfloat16 keeps the sign and exponent exactly and truncates the mantissa from 23 bits to 7, giving 1 + 8 + 7 = 16 bits. Nothing about the exponent field changes — same width, same bias of 127, same encoding of infinities and NaNs.
Converting FP32 → BF16 is therefore close to a no-op: take the top 16 bits of the 32-bit word (rounding on the boundary) and you are done. No re-scaling of the exponent, and no risk that an FP32-representable value overflows or underflows in BF16, because both formats reach the same magnitudes. Contrast float16 (FP16), which spends its bits as 1 + 5 + 10 — fewer exponent bits, a much narrower range. BF16 made the opposite bet, and for training it is the right one.
The bit layout, and a worked encoding
Here are the three formats side by side, and one number pushed through the BF16 encoder by hand.
format sign exponent mantissa total
FP32 1 8 23 32 bits
BF16 1 8 7 16 bits
FP16 1 5 10 16 bits
encode π = 3.14159 into BF16
3.14159 = 1.570796 × 2^1 (normalize)
sign = 0
exponent = 1 + 127 = 128 = 1000 0000
mantissa = .570796 → 7 bits = 100 1001
value = 1.5703125 × 2^1 = 3.140625
error = 3.14159 - 3.140625 ≈ 0.001 (≈ 0.03%)The stored value is 0 10000000 1001001. The significand is always 1.fffffff in binary — the leading 1 is implicit — so 7 stored bits buy 8 bits of significand. Because they are the top 7 fraction bits of the value FP32 would store, BF16 rounding is just ‘keep the high bits, round on bit 8’ — why hardware casts between the formats almost for free.
Range math: why 8 exponent bits reach FP32’s ceiling
Dynamic range is set entirely by the exponent field, and BF16 shares FP32’s. With an 8-bit exponent and bias 127, the stored exponent runs 0–255, with 0 and 255 reserved (subnormals/zero and infinity/NaN). So the usable normalized exponents are 1–254, i.e. actual exponents -126 to +127.
max normal = (2 - 2^-7) × 2^127 ≈ 3.39 × 10^38
min normal = 1.0 × 2^-126 ≈ 1.18 × 10^-38
min subnormal = 2^-126 × 2^-7 = 2^-133 ≈ 9.2 × 10^-40Those are, to the digit, FP32’s limits — it differs only in the (2 - 2^-23) factor at the top, a rounding hair. FP16, with 5 exponent bits, tops out at just 65504 and bottoms out near 6 × 10^-5. That gap — 10^38 versus 10^4 at the top, 10^-38 versus 10^-5 at the bottom — is why BF16 is comfortable where FP16 is fragile.
Why matching FP32’s range makes BF16 drop-in
Training a transformer produces values across an enormous span: activations of order 1 and — the dangerous end — gradients of 10^-6 or smaller for deep layers early on. In FP16, anything below ≈ 6 × 10^-5 flushes to zero, and a gradient that underflows to zero is a weight that never learns. FP16 survives this only with loss scaling: multiply the loss by a large factor (say 2^15) before the backward pass to lift small gradients into the representable band, then divide it back out before the optimizer step — usually with dynamic tuning to dodge the overflow the trick invites.
BF16 deletes this entire mechanism. Because its smallest normal is ~10^-38, gradients that would vanish in FP16 are represented directly: no scale to pick, no overflow to guard. You cast to BF16 and train. That operational simplicity, more than raw speed, is why BF16 became the default.
The precision cost: 7 mantissa bits
Range is free; precision is what you paid. With 8 significand bits, BF16’s machine epsilon — the gap between 1.0 and the next representable value — is 2^-7 = 0.0078125, roughly 2–3 significant decimal digits. FP16’s 11 significand bits give epsilon 2^-10 ≈ 0.00098 (3–4 digits), and FP32’s 2^-23 ≈ 1.2 × 10^-7 gives 7.
Crucially, floating-point precision is relative: the spacing between BF16 values scales with magnitude. In [1, 2) the step is 2^-7; in [1024, 2048) it is 2^3 = 8. So the fractional error of any single rounded value stays near 2^-8 ≈ 0.4% everywhere — harmless for one number. The trouble comes when you add millions of them, or add a number far smaller than the one you add it to — where the missing 16 mantissa bits send the bill.
Where it bites: large accumulations
Consider a dot product or a reduction over thousands of terms — the heart of every matmul and gradient sum. As a running total grows, its BF16 spacing grows with it. Once the accumulator reaches 1024, the representable gap is 8, so any term smaller than 4 rounds away to nothing. This is swamping: small contributions are silently dropped because they land below the accumulator’s current resolution.
Sum a million values of ~1.0 in pure BF16 and the total stalls far short of 1,000,000, because past a point each new +1 falls below the running sum’s rounding step and vanishes. The fix is not more mantissa bits but a wider accumulator: keep the operands in BF16, let the running sum live in FP32 — what tensor-core hardware does, and the key to accurate low-precision matmul.
Where it bites: small weight updates
The second failure mode is the optimizer step. A weight update is w ← w - η · g, and late in training η · g is routinely 10^3–10^6 times smaller than w. Watch what BF16 does to it:
w = 1.0
η · g = 0.002 (the update)
BF16 step near 1.0 = 2^-7 = 0.0078125
half-step = 0.00390625
0.002 < 0.00390625 → round-to-nearest gives 1.0
w (after) = 1.0 (the update DISAPPEARED)Because 0.002 is less than half of BF16’s step at 1.0, 1.0 + 0.002 rounds straight back to 1.0. Do this every step and the weight freezes — not because the gradient is zero, but because the update is forever smaller than the storage resolution. Thousands of real nudges accumulate to nothing, and a model that should keep improving flatlines.
The fixes: FP32 master weights and stochastic rounding
The standard cure is FP32 master weights. The optimizer keeps a full-precision copy of every parameter; the update w - η · g is applied to that FP32 copy, where 0.002 is comfortably representable and accumulates faithfully over many steps. A BF16 copy is cast off the master only for the forward and backward matmuls. Small updates land in the master and are never lost; the BF16 weights are a fast, disposable shadow.
The lighter-weight alternative is stochastic rounding: round probabilistically in proportion to distance instead of always to nearest. Here 1.0 + 0.002 rounds up to 1.0078125 with probability 0.002 / 0.0078125 ≈ 0.26 and stays at 1.0 otherwise. Any single step is wrong, but the expected value is exact, so the update survives in aggregate — letting you keep weights in BF16 without an FP32 master, at the cost of some added variance.