What emergent means here

The working definition, from Wei et al. (2022), is deliberately narrow: an ability is emergent if it is not present in smaller models but is present in larger ones. The operational test is unpredictability — you could not have forecast the large-model behavior by extrapolating the small-model trend. Below some scale the task performance hovers near random guessing regardless of how carefully you tune; above it, performance climbs well beyond chance.

This is a stronger claim than ‘bigger models are better.’ Many capabilities improve smoothly and predictably with scale — perplexity, for instance, follows a clean power law. Emergence names the opposite pattern: a flat line at chance, then a sharp rise. The word borrows from complex systems, where qualitatively new behavior appears at scale that none of the parts exhibit alone. It is descriptive, not mechanistic — it labels a shape in the data, and leaves open why the shape occurs.

Advertisement

The phenomenology: a flat line, then a jump

Plot task accuracy on the vertical axis and training compute (or parameter count) on a log horizontal axis, and an emergent ability traces a distinctive curve. For several orders of magnitude the line is flat and pinned near the random baseline: a 100M-parameter model and a 1B-parameter model are equally hopeless at the task. Then, over a relatively narrow band of scale, accuracy lifts off and keeps climbing.

The visual signature is close to a phase transition — water does not warm gradually into steam, it stays liquid and then boils. Across dozens of tasks surveyed in the original work, this same qualitative shape recurs: long dormancy, then take-off. Crucially the location of the jump differs by task — some abilities emerge at modest scale, others only in the largest models — but the shape is shared. That recurrence is what made emergence feel like a real phenomenon rather than a handful of anecdotes.

Advertisement

Canonical example: multi-digit arithmetic

The cleanest illustration is arithmetic. Ask a model to compute a 3-digit addition like 489 + 736 or a 2-digit multiplication, few-shot. Small models get essentially none right — their accuracy is indistinguishable from chance across the whole range up to a point. Then, in models around the 10^22–10^23 training-FLOP range, accuracy on 3-digit addition jumps from near zero to a substantial fraction correct.

What makes this compelling is that arithmetic is checkable — there is exactly one right answer, no partial credit, no grader subjectivity. The model is not memorizing; the test set contains number combinations it never saw verbatim. Somewhere in scaling up, the model acquires an internal procedure that carries digits and aligns place values well enough to get exact answers. It remains brittle — extend to 10-digit operands and accuracy collapses again — but the qualitative appearance of the ability is unmistakable.

Canonical example: unscrambling and transliteration

Two other favorites from the original survey are word manipulation tasks. In word unscrambling, the model is given the letters of a word in shuffled order — say l a p p e — and must recover apple. In the IPA transliteration task, it maps between ordinary spelling and phonetic notation. Both stay flat at near-zero success for small models and then emerge at scale.

These tasks matter because they are not knowledge lookups; they demand manipulating symbols according to a rule the model has to infer from a few examples. A small model that clearly ‘knows’ the word apple in its vocabulary still cannot reliably reassemble it from scrambled letters until it is large enough. That gap — possessing the pieces yet being unable to compose them — is exactly the flavor of capability that scale seems to unlock, and it recurs across many structured tasks.

In-context learning, the master ability

The most consequential emergent ability is in-context learning (ICL): the model learns a task from examples placed in the prompt, at inference time, with no gradient updates. You show it a few input→output pairs and a fresh input, and it continues the pattern. Nothing about the weights changes — the ‘learning’ happens entirely within the forward pass, conditioned on the context.

ICL is really the substrate beneath the other examples: arithmetic, unscrambling, and most benchmark tasks are posed as few-shot prompts, so the model must first grasp what task it is being asked to do from the demonstrations, then perform it. Small models largely ignore the demonstrations — adding more examples barely helps. Larger models increasingly exploit them, and the benefit of extra shots grows with scale. That shift, from ‘examples are noise’ to ‘examples are instructions,’ is itself one of the signature emergent transitions.

The GPT-3 few-shot phenomenology

GPT-3 (Brown et al., 2020) is where this became impossible to ignore. Its paper was subtitled ‘Language Models are Few-Shot Learners,’ and the central demonstration was that a 175B-parameter model, given a task description and a handful of examples in the prompt, could perform translation, question answering, cloze tasks, and arithmetic without any fine-tuning.

The paper distinguished three regimes by how many demonstrations the prompt carries: zero-shot (task description only), one-shot (a single example), and few-shot (typically 10–100 examples). The headline result was the gap between them widening with model size: for small models, few-shot barely beats zero-shot; for the 175B model, the few-shot boost is large and sometimes rivals fine-tuned systems. In other words the capacity to benefit from in-context examples was itself scaling up — the clearest early evidence that new behavior arrives with size.

Chain-of-thought as an emergent technique

A striking later finding is that some prompting techniques are themselves emergent. Chain-of-thought prompting — asking the model to write out intermediate reasoning steps before the final answer — actively hurts small models: their meandering rationales lead to worse answers than replying directly. Only above a certain scale does asking for reasoning steps start to help, and then it helps dramatically on multi-step arithmetic and word problems.

This is a useful sharpening of the concept. Emergence is not only about which tasks a model can do; it is also about which methods of eliciting behavior begin to work. A tool that is counterproductive at one scale becomes a major lever at the next. It warns against concluding a capability is absent from large models just because a technique that unlocks it does nothing on small ones — the technique and the scale have to match.

The scale thresholds, in numbers

Three quantities move together when people say ‘scale’: parameters N, training tokens D, and total compute C ≈ 6 · N · D FLOPs. Emergence is most cleanly read against compute, because it bundles the other two. The rough bands observed for several classic tasks:

TaskApprox. emergence scale
2-digit arithmetic, simple ICL~10^22 FLOPs / ~10B params
3-digit addition, word unscramble~10^23 FLOPs / tens of B params
Chain-of-thought on math word problems~10^23–10^24 FLOPs / ~100B params

Treat these as order-of-magnitude landmarks, not constants. The exact threshold for any task depends on data quality, architecture, and how the task is scored. But the qualitative message is robust: the interesting jumps cluster in the 10^22–10^24 FLOP range, which is why that band defined the frontier of capability for the GPT-3 generation of models.