The claim: abilities that switch on at scale

The canonical statement (Wei et al., 2022) is that an ability is emergent if it is not present in smaller models but is present in larger ones — and, crucially, that its appearance is not predictable by extrapolating the performance of the small models. On tasks like multi-step arithmetic, word unscrambling, or certain reasoning benchmarks, accuracy sits at chance across several orders of magnitude of scale and then shoots up once the model crosses some parameter or compute threshold N*.

Written as a curve, the claim is that capability C(N) against log-scale N looks like a step function: C(N) ≈ 0 for N < N*, then a sharp rise. If true, this is strategically enormous — it means you cannot see a capability coming, so you must simply build bigger and hope. That ‘you cannot see it coming’ part is exactly what the critical literature challenges.

Advertisement

The mathematical shape of an emergence curve

Two quantities are in play, and conflating them is the whole trap. The first is the model’s underlying quality, best captured by the smooth, predictable pre-training loss — per-token cross-entropy, which falls with scale as a power law L(N) ≈ L_∞ + (N_c / N)^α. This curve has no cliffs; it is the well-behaved object scaling laws describe.

The second is the downstream metric we report on a task — exact-match accuracy, multiple-choice accuracy, BLEU. That metric is some function M of the underlying quality: score(N) = M(L(N)). If M is smooth and roughly linear, score(N) inherits the smooth shape of L(N). But if M is sharply nonlinear or discontinuous — near-zero over a wide range of L and then rising steeply — then even a perfectly smooth L(N) produces a score(N) with a cliff. The emergence is in M, not necessarily in the model.

Advertisement

The mirage hypothesis: the ruler, not the model

This is the core of the critical position, argued most sharply by Schaeffer, Miranda, and Koyejo (2023) in ‘Are Emergent Abilities of Large Language Models a Mirage?’ Their claim is not that scaling does nothing — it is that the sharpness and unpredictability that make an ability look emergent are, in many documented cases, an artifact of the researcher’s choice of metric rather than a property of the model.

The argument has a testable structure. If emergence is a metric artifact, then (1) apparent emergence should cluster on discontinuous metrics like exact-match and multiple-choice accuracy; (2) swapping to a linear, continuous metric on the very same model checkpoints should turn the cliff into a smooth ramp; and (3) increasing the resolution of the test set (more examples, finer scoring) should reveal steady sub-threshold progress that a coarse metric rounds to zero. All three predictions can be checked against existing model families — and they largely hold.

Worked example: how exact-match manufactures a cliff

Take a task that requires emitting a sequence of n tokens all correctly — say a 5-digit arithmetic answer, scored by exact string match. Suppose per-token accuracy p(N) improves smoothly with scale. Because exact-match demands every token right, the reported score is p^n (assuming roughly independent token errors). That single exponent is enough to fabricate a cliff:

per-token pexact-match p^5looks like…
0.300.002chance / ‘no ability’
0.500.031still near zero
0.700.168‘starting to emerge’
0.900.590‘sudden’ capability
0.950.774‘mastered’

The left column marches up in even steps; the right column stays pinned near zero and then leaps. Nothing discontinuous happened to the model — the (·)^5 did all the theatrical work. Longer target sequences (larger n) make the cliff look sharper still. Treat p^n as intuition rather than gospel: real token errors are correlated, not independent, which can make the true curve gentler or sharper — but the qualitative lesson, that a harsh nonlinear metric fabricates sharpness, holds regardless.

Discontinuous metrics versus linear metrics

The distinction that matters is whether a metric gives partial credit. Discontinuous metrics do not: exact-match, multiple-choice accuracy (an argmax that is right or wrong), and pass/fail unit tests all collapse a continuum of model quality into a hard threshold. A model that is ‘almost right’ scores identically to one that is hopeless, so improvement stays invisible until it crosses the finish line all at once.

Linear or continuous metrics reward getting closer: per-token accuracy, token edit distance, log-likelihood of the correct answer, Brier score. On these, the smooth power-law improvement of the model shows through as a smooth improvement in score. The empirical punchline from the mirage work is that when you re-score the exact same checkpoints on a linear metric, the emergent cliff usually flattens into an unremarkable ramp — and, symmetrically, that you can induce apparent emergence on a boringly smooth task simply by re-scoring it with a harsh nonlinear metric. The phenomenon tracks the ruler, not the task.

What the critique does not claim

It is easy to over-read the mirage result into ‘emergence is fake, scaling buys nothing new.’ That is not the argument, and the distinction is worth guarding. Scaling demonstrably yields models that solve tasks smaller ones cannot; the smooth loss curve is still going somewhere, and a capability that is at 1% and one at 80% are meaningfully different even if the path between them was continuous.

What the critique narrows is the strong, spooky version of the claim: that capabilities appear discontinuously and unpredictably. If the underlying quality improves smoothly and predictably, then with the right metric you can forecast when a task will cross a usefulness threshold — the surprise was an artifact of measurement, not a fundamental unknowability. Some abilities may still exhibit genuine phase-transition dynamics; the point is that the burden of proof is on showing the sharpness survives a continuous metric.

Grokking: a real sharp transition, in training not scale

The most interesting genuinely-abrupt phenomenon is grokking (Power et al., 2022): on small algorithmic tasks, a network first memorizes the training set (train accuracy → 100%, test accuracy at chance), then, after many more steps of seemingly pointless training, test accuracy jumps suddenly to near-perfect. It is a real discontinuity in generalization — but note the axis: it happens over training time at fixed model size, not over scale.

Grokking is often driven by regularization (weight decay) slowly pushing the network from a memorizing solution to a structured, generalizing one — mechanistic studies find a clean circuit crystallizing internally. So grokking is evidence that sharp transitions can be real, located in weight space, not just a scoring artifact. It is the honest counterweight to the mirage view: some cliffs are in the measurement, but at least one well-studied cliff is genuinely in the model’s internal representation.