The question that reorganized an industry
Every large-model project starts from the same fork: you have a fixed compute budget C, and you must split it between a bigger model (more parameters N) and more training data (more tokens D). Because training cost is roughly C ≈ 6ND, the two trade off directly — every parameter you add is tokens you cannot afford, and vice versa.
What makes this more than an engineering detail is that the answer sets the entire shape of a model program: how many GPUs, how large a dataset to collect and clean, how much the finished model will cost to serve. Get the split wrong and you either burn compute on parameters you cannot feed, or you starve a capable architecture of data. For years the field answered ‘favor parameters,’ and built accordingly. Chinchilla’s lasting contribution was less a single number than a correction to how the whole question was being answered.
The era of the giant under-trained model
The pre-Chinchilla generation is striking in hindsight. GPT-3 was 175B parameters trained on ~300B tokens — a data-to-parameter ratio under 2. Gopher pushed to 280B parameters on the same ~300B tokens (ratio ~1). Megatron-Turing NLG reached 530B parameters on ~270B tokens. The consistent pattern: parameters climbed into the hundreds of billions while token counts barely moved.
This was not carelessness; it followed the best guidance available. The Kaplan et al. (2020) scaling study had concluded that parameters should grow far faster than data, so pouring budget into size looked optimal. The result was a fleet of enormous models that, we now know, were badly under-trained — they had far more capacity than the data they saw could ever fill. They were expensive to train, expensive to run, and quietly leaving performance on the table that a smaller, better-fed model would have captured for free.
Kaplan and Hoffmann: two answers to one question
The Chinchilla team re-ran the scaling study with a crucial fix — matching each training run’s learning-rate schedule to its actual token budget — and the conclusion flipped. Where Kaplan had parameters dominating (N ∝ C^0.73, data almost an afterthought), Hoffmann found parameters and data scaling almost identically: N ∝ C^0.46 and D ∝ C^0.54, both close to √C.
The practical gap between those two answers is enormous. Under Kaplan, a 10× compute increase buys mostly a bigger model. Under Chinchilla, it buys a model and a dataset each grown by about √10 ≈ 3.2× — same budget, radically different machine. The correction had a story of its own (a schedule artifact that unfairly penalized data-rich runs), but the headline for practitioners was blunt: stop building giant models on tiny datasets. Feed them.
The head-to-head that made the point
Chinchilla did not just argue the theory; it built the counter-example. The same team trained a model at exactly the compute budget they had spent on Gopher — but allocated the Chinchilla way: 70B parameters on 1.4T tokens (ratio 20) instead of Gopher’s 280B on 300B tokens (ratio ~1).
The 70B Chinchilla, four times smaller than Gopher, outperformed it across the great majority of benchmarks — language modeling, reading comprehension, common-sense reasoning, MMLU. That result is what turned a scaling paper into a strategic reset. It was not a marginal efficiency tweak; it was a demonstration that a quarter of the parameters, spent on the same compute but more data, produced a better model. And because it was smaller, Chinchilla was also cheaper and faster to run — a preview of the inference argument that would soon dominate the conversation.
What compute-optimal quietly leaves out
Here is the fine print that reshapes everything downstream: Chinchilla-optimal minimizes training loss for a fixed training budget. It is silent on the cost of using the model afterward. The optimization pretends training is the only bill you will ever pay.
For a research artifact trained once and probed a few times, that assumption is fine. For a product, it is unrealistic. A deployed model may serve trillions of tokens over its lifetime, and each costs about 2N FLOPs at inference — a cost that scales with the parameter count you chose. Two models of equal quality but different sizes are not equally good products: the smaller one is cheaper on every request, forever. Once you count inference, the question is no longer ‘what minimizes training loss?’ but ‘what minimizes total cost for a given quality?’ — and the answer moves.
Over-training: trading training FLOPs for inference savings
The move the industry made after Chinchilla was to deliberately train past the 20:1 point. If a model will be served heavily, you pick a smaller N than compute-optimal and pour in more data than 20:1 to recover the quality. You spend extra training compute up front to buy a permanently cheaper model to run.
Llama is the canonical example. Llama-2 7B saw ~2T tokens — a ratio near 280:1, more than ten times past Chinchilla-optimal. Llama-3 8B went further, to ~15T tokens. By Chinchilla’s training-only accounting these models are ‘over-trained’ and technically sub-optimal. By the accounting that includes inference, they are exactly right: the extra training pass is a one-time cost, amortized across billions of cheap forward passes on a small model that fits comfortably on modest hardware. The over-training premium is small; the lifetime inference saving is not.
The inference-optimal frontier
Formally, over-training just means optimizing a different objective. Instead of minimizing training loss subject to C_train, you minimize total cost — training plus expected inference — for a target quality:
C_total ≈ 6·N·D + 2·N·D_servewhere D_serve is the total tokens you expect to process at inference over the model’s life. When D_serve is large — a popular product — the second term dominates, and it depends on N but not on training D. So the optimizer pushes N down and lets training D rise to hold quality steady. The bigger your expected serving volume, the smaller and more heavily-trained your model should be. Chinchilla’s 20:1 is simply the special case where D_serve is zero — the frontier you follow when nobody will ever run the model.
The data wall
Over-training and equal-scaling both lean on the same silent assumption: that there is always more high-quality data to reach for. At frontier scale that assumption starts to break. The stock of high-quality public text on the web — well-written, deduplicated, non-spam — has been estimated in the low tens of trillions of tokens. Frontier training runs are now measured in tens of trillions. The curves are crossing.
This is the data wall: a point where the recipe calls for more clean tokens than exist. When D_optimal exceeds your usable corpus, the elegant scaling math quietly stops applying — there is no fresh data left to buy quality with. Compute keeps getting cheaper; unique high-quality human text does not. Increasingly the binding constraint on the best models is not FLOPs but tokens, which inverts the premise Chinchilla was built on.