The one-line idea: FP32 with the tail chopped off

A float32 lays out its 32 bits as 1 sign, 8 exponent, 23 mantissa. bfloat16 keeps the sign and exponent exactly and truncates the mantissa from 23 bits to 7, giving 1 + 8 + 7 = 16 bits. Nothing about the exponent field changes — same width, same bias of 127, same encoding of infinities and NaNs.

Converting FP32 → BF16 is therefore close to a no-op: take the top 16 bits of the 32-bit word (rounding on the boundary) and you are done. No re-scaling of the exponent, and no risk that an FP32-representable value overflows or underflows in BF16, because both formats reach the same magnitudes. Contrast float16 (FP16), which spends its bits as 1 + 5 + 10 — fewer exponent bits, a much narrower range. BF16 made the opposite bet, and for training it is the right one.

Advertisement

The bit layout, and a worked encoding

Here are the three formats side by side, and one number pushed through the BF16 encoder by hand.

format   sign  exponent  mantissa   total
FP32       1       8         23       32 bits
BF16       1       8          7       16 bits
FP16       1       5         10       16 bits

encode π = 3.14159 into BF16
  3.14159  = 1.570796 × 2^1      (normalize)
  sign     = 0
  exponent = 1 + 127 = 128  = 1000 0000
  mantissa = .570796 → 7 bits = 100 1001
  value    = 1.5703125 × 2^1 = 3.140625
  error    = 3.14159 - 3.140625 ≈ 0.001  (≈ 0.03%)

The stored value is 0 10000000 1001001. The significand is always 1.fffffff in binary — the leading 1 is implicit — so 7 stored bits buy 8 bits of significand. Because they are the top 7 fraction bits of the value FP32 would store, BF16 rounding is just ‘keep the high bits, round on bit 8’ — why hardware casts between the formats almost for free.

Advertisement

Range math: why 8 exponent bits reach FP32’s ceiling

Dynamic range is set entirely by the exponent field, and BF16 shares FP32’s. With an 8-bit exponent and bias 127, the stored exponent runs 0–255, with 0 and 255 reserved (subnormals/zero and infinity/NaN). So the usable normalized exponents are 1–254, i.e. actual exponents -126 to +127.

max normal = (2 - 2^-7) × 2^127   ≈ 3.39 × 10^38
min normal = 1.0        × 2^-126  ≈ 1.18 × 10^-38
min subnormal = 2^-126 × 2^-7 = 2^-133 ≈ 9.2 × 10^-40

Those are, to the digit, FP32’s limits — it differs only in the (2 - 2^-23) factor at the top, a rounding hair. FP16, with 5 exponent bits, tops out at just 65504 and bottoms out near 6 × 10^-5. That gap — 10^38 versus 10^4 at the top, 10^-38 versus 10^-5 at the bottom — is why BF16 is comfortable where FP16 is fragile.

Why matching FP32’s range makes BF16 drop-in

Training a transformer produces values across an enormous span: activations of order 1 and — the dangerous end — gradients of 10^-6 or smaller for deep layers early on. In FP16, anything below ≈ 6 × 10^-5 flushes to zero, and a gradient that underflows to zero is a weight that never learns. FP16 survives this only with loss scaling: multiply the loss by a large factor (say 2^15) before the backward pass to lift small gradients into the representable band, then divide it back out before the optimizer step — usually with dynamic tuning to dodge the overflow the trick invites.

BF16 deletes this entire mechanism. Because its smallest normal is ~10^-38, gradients that would vanish in FP16 are represented directly: no scale to pick, no overflow to guard. You cast to BF16 and train. That operational simplicity, more than raw speed, is why BF16 became the default.

The precision cost: 7 mantissa bits

Range is free; precision is what you paid. With 8 significand bits, BF16’s machine epsilon — the gap between 1.0 and the next representable value — is 2^-7 = 0.0078125, roughly 2–3 significant decimal digits. FP16’s 11 significand bits give epsilon 2^-10 ≈ 0.00098 (3–4 digits), and FP32’s 2^-23 ≈ 1.2 × 10^-7 gives 7.

Crucially, floating-point precision is relative: the spacing between BF16 values scales with magnitude. In [1, 2) the step is 2^-7; in [1024, 2048) it is 2^3 = 8. So the fractional error of any single rounded value stays near 2^-8 ≈ 0.4% everywhere — harmless for one number. The trouble comes when you add millions of them, or add a number far smaller than the one you add it to — where the missing 16 mantissa bits send the bill.

Where it bites: large accumulations

Consider a dot product or a reduction over thousands of terms — the heart of every matmul and gradient sum. As a running total grows, its BF16 spacing grows with it. Once the accumulator reaches 1024, the representable gap is 8, so any term smaller than 4 rounds away to nothing. This is swamping: small contributions are silently dropped because they land below the accumulator’s current resolution.

Sum a million values of ~1.0 in pure BF16 and the total stalls far short of 1,000,000, because past a point each new +1 falls below the running sum’s rounding step and vanishes. The fix is not more mantissa bits but a wider accumulator: keep the operands in BF16, let the running sum live in FP32 — what tensor-core hardware does, and the key to accurate low-precision matmul.

Where it bites: small weight updates

The second failure mode is the optimizer step. A weight update is w ← w - η · g, and late in training η · g is routinely 10^3–10^6 times smaller than w. Watch what BF16 does to it:

w            = 1.0
η · g        = 0.002        (the update)
BF16 step near 1.0 = 2^-7 = 0.0078125
half-step          = 0.00390625
0.002 < 0.00390625  → round-to-nearest gives 1.0
w (after)    = 1.0        (the update DISAPPEARED)

Because 0.002 is less than half of BF16’s step at 1.0, 1.0 + 0.002 rounds straight back to 1.0. Do this every step and the weight freezes — not because the gradient is zero, but because the update is forever smaller than the storage resolution. Thousands of real nudges accumulate to nothing, and a model that should keep improving flatlines.

The fixes: FP32 master weights and stochastic rounding

The standard cure is FP32 master weights. The optimizer keeps a full-precision copy of every parameter; the update w - η · g is applied to that FP32 copy, where 0.002 is comfortably representable and accumulates faithfully over many steps. A BF16 copy is cast off the master only for the forward and backward matmuls. Small updates land in the master and are never lost; the BF16 weights are a fast, disposable shadow.

The lighter-weight alternative is stochastic rounding: round probabilistically in proportion to distance instead of always to nearest. Here 1.0 + 0.002 rounds up to 1.0078125 with probability 0.002 / 0.0078125 ≈ 0.26 and stays at 1.0 otherwise. Any single step is wrong, but the expected value is exact, so the update survives in aggregate — letting you keep weights in BF16 without an FP32 master, at the cost of some added variance.