Scaling laws as a budgeting tool
Strip away the curve-fitting and a scaling law answers one planning question: given a fixed training budget, what allocation of that budget between model size and data minimises loss? Both knobs cost compute. A bigger model does more arithmetic per token; more tokens means more steps. You cannot maximise both, so there is a genuine optimum, and it is that optimum the Chinchilla-style fits locate.
Two identities carry almost all the practical weight. Training a dense transformer costs roughly C_train ≈ 6 · N · D FLOPs — two FLOPs per parameter forward, about four backward. Serving it costs roughly C_token ≈ 2 · N FLOPs per token, forward pass only. Notice the asymmetry: training cost scales with N × D and is paid once; inference cost scales with N alone and is paid every single time anyone uses the model. D is free at serving time. That single observation is the whole architecture story.
The compute-optimal rule of thumb
The headline result of the Chinchilla work is easy to state and easy to apply: at the compute-optimal point, N and D should scale together, roughly in the ratio 20 training tokens per parameter. (Why the exponents come out near-equal belongs to the derivation of the loss fit; here we simply use the answer.)
Combine D = 20N with C = 6ND and you get a closed-form sizing rule you can run in your head:
C = 6 · N · (20N) = 120 · N^2
N* = sqrt(C / 120) D* = 20 · N*
worked: C = 1e24 FLOPs
N* = sqrt(1e24 / 120) = sqrt(8.33e21) ≈ 9.1e10 (~91B params)
D* = 20 × 9.1e10 ≈ 1.8e12 (~1.8T tokens)
check: 6 × 9.1e10 × 1.8e12 ≈ 1.0e24 ✓Run it backwards too. If someone hands you a 7B model and says it was trained on 140B tokens, that is exactly 20 tokens per parameter — textbook compute-optimal, and a strong hint that the team optimised for training cost rather than serving cost.
When compute-optimal is genuinely the right target
Before attacking the rule, be clear about when it is correct, because it often is. Compute-optimal is the right target whenever training dominates lifetime cost — which is to say, whenever the model will not be run very much.
That covers more of the field than people assume: research runs and ablations, where the artefact exists to produce a number in a table; capability probes, where nobody will ever serve the checkpoint; internal one-off evaluations; and frontier races where being first is worth more than the serving margin. In all of these the model is trained once and inferenced comparatively little, so minimising L subject to C_train is exactly right.
It is also the correct starting point even when you intend to deviate: it tells you where the frontier sits, so you can move along it deliberately. Deviating without knowing the reference point is a guess.
The inference frontier: when serving flips the objective
Now the production case. You are not buying a loss number; you are buying a product that will answer requests for a year or more. The objective is no longer minimise loss given training compute. It is minimise total compute (training plus all inference) subject to hitting a quality bar. That is a different optimisation and it has a different answer.
Because inference cost scales with N and not with D, the way to cut serving cost is to shrink the model — and the way to keep quality while shrinking it is to feed it far more tokens than Chinchilla prescribes. This is overtraining: deliberately moving off the compute-optimal point to a smaller N with a much larger D. You spend more training compute than optimal to reach the same quality, and you get a permanently cheaper model out of it. Every widely deployed 7B and 8B workhorse of the last few years sits far out in this regime — not because their authors miscalculated, but because they were pricing the whole lifetime.
The crossover: training FLOPs versus lifetime inference FLOPs
The trade has an exact break-even. Compare two roughly iso-quality configurations — a compute-optimal model (N_1, D_1) and a smaller overtrained one (N_2, D_2) with N_2 < N_1 and D_2 > D_1 — over a lifetime of T served tokens. Set total costs equal and solve:
6·N_1·D_1 + 2·N_1·T = 6·N_2·D_2 + 2·N_2·T
T* = 3 · (N_2·D_2 − N_1·D_1) / (N_1 − N_2)
worked: A = 30B params / 600B tokens (20 tok/param, compute-optimal)
B = 8B params / 6T tokens (750 tok/param, overtrained)
train A = 6 × 3.0e10 × 6.0e11 = 1.08e23 FLOPs
train B = 6 × 8.0e9 × 6.0e12 = 2.88e23 FLOPs (2.7× more)
per served token: A = 2N = 6.0e10 B = 2N = 1.6e10 (3.75× cheaper)
T* = (2.88e23 − 1.08e23) / (6.0e10 − 1.6e10) ≈ 4.1e12 tokensSo B repays its extra training compute after about 4 trillion served tokens — on the order of a billion or two real requests. Past that point every token is 3.75× cheaper, forever. Below it, you overpaid.
How far past Chinchilla to go
‘Overtrain’ is a direction, not a destination. The useful parameter is the token-per-parameter multiple, and shipped models span two orders of magnitude on it:
| Regime | Tokens / param | Who lives here |
|---|---|---|
| Compute-optimal | ~20 | Research runs, one-off checkpoints |
| Mildly overtrained | ~100–300 | Mid-size served models |
| Heavily overtrained | ~1000–2000 | Modern 7B–8B workhorses |
| Extreme | 2000+ | Sub-3B edge and on-device models |
The stopping rule follows from the shapes of the two curves. Serving savings are linear in N: halve the model, halve the cost per token. Quality recovery is not — each further increment of loss you try to claw back costs disproportionately more data. So the multiple where you stop is the one where the next doubling of D costs more training compute than the shrink saves you over your actual projected T. If you find yourself needing a 4× data increase to hold quality, the honest answer is usually to take the larger model instead.