GPT-2: scale turns a language model into a task solver

The 2019 GPT-2 work (Radford et al., Language Models are Unsupervised Multitask Learners) made the first loud scaling argument. A single transformer trained only to predict the next token on a large, diverse web corpus could, at sufficient size, perform reading comprehension, translation, summarization, and question answering zero-shot — with no task-specific fine-tuning, just a prompt. Smaller versions of the same model mostly could not.

The framing was decisive: language modeling is not a narrow objective but a compression of the tasks latent in text, and scale is what unlocks them. Performance on these downstream tasks rose smoothly and log-linearly with model size across the four GPT-2 variants, foreshadowing the formal laws. The lesson that shaped everything after: you do not need a bespoke model per task if one large enough general model absorbs the tasks for free. Capability, in this view, is something you grow rather than something you hand-engineer.

Advertisement

The compute-efficient frontier

One of OpenAI’s most useful reframings is the compute-efficient frontier. Plot the training loss curve of every model in a sweep against compute spent, and the lower envelope of all those curves — the best loss achievable for each compute budget — is itself a clean power law. Any single model rides that envelope for a while, then peels off it as it saturates and its curve flattens.

The practical reading is counterintuitive: training a model to convergence is wasteful. Long before a given model bottoms out, a larger model has already passed it on the envelope for the same compute. So the compute-optimal move is to stop each model early, while it is still improving fast, and pour the saved compute into a bigger one. This is a statement about training dynamics, not the final N-vs-D allocation — it says the best use of the next FLOP is usually a step on a larger model, not another step on the current one. It is why frontier runs look under-trained by classical convergence standards.

Advertisement

Larger models are more sample-efficient

The frontier picture rests on a finding that surprised many people: bigger models are more sample-efficient, not less. A larger model reaches any given loss in fewer optimization steps and after seeing fewer tokens than a smaller one. Extra parameters do not just raise the ceiling; they make each gradient step more productive.

The intuition runs against everyday experience with over-parameterized models and overfitting, but in the large-data pretraining regime it holds robustly: a 10B model does not merely end up better than a 1B model — it gets to the 1B model’s final loss much sooner. Combined with the near-negligible cost of over-parameterization in this regime, this is what makes ‘train big, stop early’ rational, and it connects to OpenAI’s work on the critical batch size (McCandlish et al.), which bounds how much of that faster learning you can convert into shorter wall-clock time by scaling batch size.

GPT-3 and emergent in-context learning

GPT-3 (Brown et al., 2020) was the program’s boldest extrapolation: a 175-billion-parameter model, roughly a 100× jump over GPT-2, staked on the prediction that the curves would keep going. They did — loss landed near where the laws said it would — but the headline was a capability the loss number did not obviously advertise. At scale the model exhibited in-context learning: shown a few examples of a task in its prompt, it adapts to that task at inference time, with no weight updates at all.

Few-shot in-context learning strengthened with scale and was weak or absent in smaller models. This is the sharpest evidence for the program’s thesis that scale produces qualitatively new behavior, not just incremental loss reduction. Smooth pretraining loss and sudden-looking downstream jumps coexist: the loss glides down its line while a benchmark it enables crosses a usefulness threshold. GPT-3 turned scaling from an internal research bet into an industry-wide one.

The same curves hold beyond text

A fair worry about the language laws was that they might be a quirk of natural language. OpenAI’s multimodal study (Henighan et al., 2020, Scaling Laws for Autoregressive Generative Modeling) tested that directly and found the power laws are universal across modalities — images, video, math problems, and image-text pairs all show the same smooth, power-law fall in loss with compute and model size.

That work also sharpened the functional form with an explicit irreducible term, L(C) = L_∞ + (C_c / C)^α_C. The L_∞ floor is the entropy of the data itself — the loss no amount of scale can remove — and the second term is the reducible part that scaling eats into. The universality across domains implied that scaling is a property of large autoregressive models in general, not a linguistic accident — which is what justified betting on scale for multimodal frontier systems.

Scaling laws for transfer

Pretraining is only worthwhile if the loss it lowers translates into skill on the tasks you actually care about. The transfer study (Hernandez et al., 2021, Scaling Laws for Transfer) quantified that bridge. It measured the effective data transferred: how much fine-tuning data a pretrained model is effectively worth, expressed as the amount of target-domain data a from-scratch model would need to match it.

The finding is that pretraining acts like a multiplier on your fine-tuning set, and the multiplier is largest exactly where you need it — the low-data regime. When target data is scarce, a large pretrained model is worth an enormous quantity of effective task data; as you accumulate more real target data, the relative bonus shrinks. This gave scaling an economic argument beyond raw loss: money spent on a bigger, better-pretrained base pays off as data you did not have to collect for every downstream task.