The exact-match trap: multiplying probabilities
Fix a task whose correct answer is a specific string of n tokens — a 12-digit product, a 6-step proof, a JSON object with n fields. Suppose the model gets each token right independently with probability p. Then the probability that every token is correct, which is what an exact-match grader rewards, is the product
EM = P(all n correct) = p · p · ... · p = p^n
log EM = n · log p (log p < 0, so EM decays)Independence is an idealization — real errors are correlated, and a model that recovers from a slip does better than the product predicts — but correlation changes the constants, not the shape. The load-bearing fact is that EM = p^n decays exponentially in the task length n. Because log p is negative, every extra required token subtracts a fixed amount from log EM. A per-token accuracy that feels excellent in isolation — 95% — is quietly fatal over a long chain, and that single exponential is the whole engine of emergence.
A worked table of p to the n
Nothing makes the point faster than tabulating p^n. Read down a column and watch a ‘good’ per-token accuracy collapse as the task gets longer; read across a row and watch a long task stay near zero until p is almost perfect.
| per-token p | n = 1 | n = 5 | n = 20 | n = 50 |
|---|---|---|---|---|
| 0.50 | 0.50 | 0.031 | 0.0000010 | ≈ 0 |
| 0.80 | 0.80 | 0.328 | 0.012 | 0.000014 |
| 0.90 | 0.90 | 0.590 | 0.122 | 0.0052 |
| 0.95 | 0.95 | 0.774 | 0.358 | 0.077 |
| 0.99 | 0.99 | 0.951 | 0.818 | 0.605 |
The n = 50 column is the story. As per-token accuracy climbs from 0.90 to 0.99 — a smooth, unremarkable tenfold reduction in error rate — exact-match success on the long task leaps from 0.5% to 61%. Nothing discontinuous happened to the model; we multiplied a number just under 1 by itself fifty times and the result is savagely sensitive to that number. The metric manufactures a cliff out of a ramp.
Why the curve looks like a cliff
To see the cliff analytically, work near p = 1, where the per-token error is e = 1 − p. Using log(1 − e) ≈ −e for small e,
EM = p^n = (1 − e)^n ≈ exp(−n · e)So exact-match is governed by the single product n · e. When the error e halves, the exponent halves, and the new exact-match score is the square root of the old one. Take a task sitting at EM = 0.01: halve its per-token error and it jumps to √0.01 = 0.10, a tenfold gain from one modest improvement. Halve it again and you are at 0.32. Rooting a small number repeatedly is a fast climb — precisely the regime a scaling model passes through as it drives e down. The steepness is not evidence of a phase change inside the network — it is the geometry of exp(−n e) as e → 0.
The sigmoid in log-compute
Now feed in what scaling laws actually give us. Empirically the per-token loss, and with it the error rate, falls as a power law in compute C: e(C) ≈ a · C^(−α) for constants a > 0 and small exponent α. Substitute into EM ≈ exp(−n e):
EM(C) ≈ exp( −n · a · C^(−α) )
let x = log C ⇒ EM ≈ exp( −n a · e^(−α x) )That double exponential is a Gompertz curve — a sigmoid when plotted against log C. At low compute C^(−α) is large, the inner term is huge, and EM ≈ 0: the ability is simply absent. As compute grows the inner term collapses toward zero and EM → 1. In between sits a narrow knee where the score sweeps from near-0 to near-1. Plot the same quantity against raw compute and you get a boring flat-then-flat shape; plot it against log compute — the axis scaling curves are always drawn on — and you get the textbook emergence S. The curve is real; the surprise is not.
Where the threshold sits, and why longer tasks emerge later
Define the emergence threshold as the compute at which the ability crosses the halfway mark, EM(C*) = 1/2. Setting the exponent equal to log 2 gives a clean closed form:
n · a · C*^(−α) = log 2
C* = ( n · a / log 2 )^(1/α) ⇒ C* ∝ n^(1/α)Two consequences follow. First, the threshold moves to the right as the task lengthens: a 50-step ability emerges at strictly more compute than a 5-step ability, scaling as n^(1/α). Second — and this is why the jumps look so alike across benchmarks — the width of the transition in log-compute is set by α alone and is essentially independent of n. Increasing n slides a rigid step-shape rightward rather than stretching it. A family of abilities of different lengths therefore appears as a series of similarly sharp steps marching along the compute axis — exactly the picture the emergence literature reports, reproduced from two constants.
Continuous surrogate metrics linearize the curve
Here is the pivot the whole subject turns on. Suppose that instead of all-or-nothing exact match we score the expected fraction of correct tokens. With the same per-token accuracy that is simply
E[fraction correct] = p = 1 − a · C^(−α)This rises smoothly and monotonically with compute — no knee, no step, just the gentle power-law approach to 1. The same is true of other continuous surrogates: per-token cross-entropy (the training loss itself), token-level edit distance, a Brier score on the answer distribution, or any partial-credit rubric. Each is roughly linear in p, so each inherits p’s smooth trajectory. The sharpness lived entirely in the nonlinear p^n product and the 0/1 threshold on top of it — not in the model, whose underlying competence was improving at a steady, predictable rate the whole time. Which metric you pick decides whether you ‘see’ an emergent jump at all; the deeper argument about what that implies is the sibling article’s.
The derivative view: locating the steepest point
How steep, and steepest where? Differentiate the Gompertz form with respect to x = log C. Writing u(x) = n a · e^(−α x) so that EM = e^(−u),
d(EM)/dx = −e^(−u) · du/dx = EM · α · u
at the knee u = log 2: slope = (1/2) · α · log 2 ≈ 0.35 · αThe slope is a product of the current score EM and the shrinking driver u; it vanishes at both ends (where either EM or u is near zero) and peaks in the middle. Its maximum magnitude is proportional to α, the scaling exponent — a small α means a gentle loss curve yet still a visibly sharp EM step, because the sharpness is amplified by the exponential, not by α being large. Crucially the peak location depends on n (through the knee) but the peak slope does not — the analytic restatement of the ‘rigid step that slides’ observation from before.