A straight line on a log-log plot
The empirical fact that starts everything: plot test loss against parameter count with both axes logarithmic, and the points fall on a straight line over several decades. A straight line in log-log coordinates is the signature of a power law, because log L = −α log N + const exponentiates to L ∝ N^−α. Kaplan et al. (2020) wrote this in the normalized form:
L(N) = (N_c / N)^α_N with α_N ≈ 0.076, N_c ≈ 8.8 × 10^13The contrast worth internalizing is with an exponential, L ∝ e^−kN, which would be a straight line on a semi-log plot and would mean each fixed increment of parameters buys the same multiplicative gain. Reality is far less generous. A power law means each fixed multiplication of parameters buys a fixed multiplicative gain: with α_N = 0.076, doubling N multiplies the loss by 2^−0.076 ≈ 0.949 — about five percent, forever. That is the whole economics of scaling in one number.
Three single-variable laws, each with a fine print
Kaplan et al. reported the same shape three times, once for each resource: parameters N, dataset size D in tokens, and compute C in floating-point operations. Each is a clean power law, with exponents around α_N ≈ 0.076, α_D ≈ 0.095, and α_C ≈ 0.050 in their parameterization.
The fine print matters more than the numbers. Each single-variable law is measured with the other resource held non-binding: L(N) is the loss of a model of size N trained to convergence with effectively unlimited data, and L(D) is the loss of a very large model trained on D tokens with early stopping. Fit one of these curves on runs where the other resource was quietly the bottleneck and you measure a blend of two effects, not a clean exponent. This is precisely the trap that makes a joint two-variable model necessary, and it is where the two famous papers part company.
The joint form and the irreducible floor
The form that has held up is additive and separable in the two reducible sources of error:
L(N, D) = E + A / N^α + B / D^βRead it as a risk decomposition. E is the irreducible term: the entropy of natural text under the true distribution, the loss a perfect predictor would still pay because language is genuinely stochastic. No amount of parameters or tokens removes it, and it is why the log-log line must eventually bend flat rather than descend forever. A/N^α is approximation error — a finite transformer cannot represent the ideal predictor, and this term measures the gap. B/D^β is estimation error — even a perfectly expressive model must infer its parameters from a finite sample. Crucially the power law lives only in the reducible part; the raw loss L is not a power law at all, which is exactly why fitting one to raw loss over a narrow range gives a misleadingly shallow exponent.
Chinchilla's fitted constants, and what they weigh
Hoffmann et al. (2022) fit that five-parameter form to over 400 training runs spanning roughly 70M to 16B parameters and 5B to 400B tokens, minimizing a Huber loss on log L for robustness against outliers. The fit:
| Constant | Value | Meaning |
|---|---|---|
E | 1.69 | Irreducible loss, nats/token |
A | 406.4 | Scale of the parameter term |
B | 410.7 | Scale of the data term |
α | 0.34 | Parameter exponent |
β | 0.28 | Data exponent |
Two things stand out. First, α > β: the parameter term decays faster in its own resource — which, counterintuitively, is exactly why the optimization below will tell you to scale the data side harder. Second, the exponents are roughly four times larger than Kaplan’s. That is not a contradiction — they are exponents of different functions. Kaplan’s 0.076 describes raw loss including the floor; 0.34 describes the reducible part after the floor is subtracted. Comparing them directly is a category error.
Why Kaplan and Chinchilla diverged
Both papers were competently executed, and both fit real data. Three methodological differences explain the gap between Kaplan’s N* ∝ C^0.73 and Chinchilla’s near-C^0.5.
The learning-rate schedule. Kaplan’s runs used a cosine decay set to a fixed horizon, then read off intermediate checkpoints as if each were a completed run. A checkpoint halfway through a cosine schedule has not annealed, and its loss is systematically worse than a run that was scheduled to end there. That penalizes long-training runs and makes data look less valuable than it is. Chinchilla matched the decay length to each run’s own token budget.
Embedding parameters. Kaplan excluded embedding and unembedding parameters from N. For small models those are a large fraction of the total, so the smallest models were credited with far fewer parameters than they had, tilting the fitted slope. Chinchilla counted all parameters. Range compounded both: a narrower ladder of model sizes leaves less leverage to separate the two exponents.
The constrained optimization
Now the central derivation. Training compute for a dense transformer is well approximated by C ≈ 6ND — roughly 2 FLOPs per parameter per token forward and 4 backward. Fix a budget C and ask which (N, D) on that hyperbola minimizes loss. Form the Lagrangian J = L(N, D) + λ(6ND − C) and set both partials to zero:
∂J/∂N = −αA·N^−(α+1) + 6λD = 0
∂J/∂D = −βB·D^−(β+1) + 6λN = 0
multiply the first by N, the second by D, and equate 6λND:
α · (A / N^α) = β · (B / D^β)That last line is the whole optimum in one equation, and it is elegant: at the compute-optimal point the two reducible loss terms are held in a fixed ratio, weighted by their own exponents. Substituting D = C/(6N) and solving for N gives the closed form.