A straight line on a log-log plot

The empirical fact that starts everything: plot test loss against parameter count with both axes logarithmic, and the points fall on a straight line over several decades. A straight line in log-log coordinates is the signature of a power law, because log L = −α log N + const exponentiates to L ∝ N^−α. Kaplan et al. (2020) wrote this in the normalized form:

L(N) = (N_c / N)^α_N     with  α_N ≈ 0.076,  N_c ≈ 8.8 × 10^13

The contrast worth internalizing is with an exponential, L ∝ e^−kN, which would be a straight line on a semi-log plot and would mean each fixed increment of parameters buys the same multiplicative gain. Reality is far less generous. A power law means each fixed multiplication of parameters buys a fixed multiplicative gain: with α_N = 0.076, doubling N multiplies the loss by 2^−0.076 ≈ 0.949 — about five percent, forever. That is the whole economics of scaling in one number.

Advertisement

Three single-variable laws, each with a fine print

Kaplan et al. reported the same shape three times, once for each resource: parameters N, dataset size D in tokens, and compute C in floating-point operations. Each is a clean power law, with exponents around α_N ≈ 0.076, α_D ≈ 0.095, and α_C ≈ 0.050 in their parameterization.

The fine print matters more than the numbers. Each single-variable law is measured with the other resource held non-binding: L(N) is the loss of a model of size N trained to convergence with effectively unlimited data, and L(D) is the loss of a very large model trained on D tokens with early stopping. Fit one of these curves on runs where the other resource was quietly the bottleneck and you measure a blend of two effects, not a clean exponent. This is precisely the trap that makes a joint two-variable model necessary, and it is where the two famous papers part company.

Advertisement

The joint form and the irreducible floor

The form that has held up is additive and separable in the two reducible sources of error:

L(N, D) = E + A / N^α + B / D^β

Read it as a risk decomposition. E is the irreducible term: the entropy of natural text under the true distribution, the loss a perfect predictor would still pay because language is genuinely stochastic. No amount of parameters or tokens removes it, and it is why the log-log line must eventually bend flat rather than descend forever. A/N^α is approximation error — a finite transformer cannot represent the ideal predictor, and this term measures the gap. B/D^β is estimation error — even a perfectly expressive model must infer its parameters from a finite sample. Crucially the power law lives only in the reducible part; the raw loss L is not a power law at all, which is exactly why fitting one to raw loss over a narrow range gives a misleadingly shallow exponent.

Chinchilla's fitted constants, and what they weigh

Hoffmann et al. (2022) fit that five-parameter form to over 400 training runs spanning roughly 70M to 16B parameters and 5B to 400B tokens, minimizing a Huber loss on log L for robustness against outliers. The fit:

ConstantValueMeaning
E1.69Irreducible loss, nats/token
A406.4Scale of the parameter term
B410.7Scale of the data term
α0.34Parameter exponent
β0.28Data exponent

Two things stand out. First, α > β: the parameter term decays faster in its own resource — which, counterintuitively, is exactly why the optimization below will tell you to scale the data side harder. Second, the exponents are roughly four times larger than Kaplan’s. That is not a contradiction — they are exponents of different functions. Kaplan’s 0.076 describes raw loss including the floor; 0.34 describes the reducible part after the floor is subtracted. Comparing them directly is a category error.

Why Kaplan and Chinchilla diverged

Both papers were competently executed, and both fit real data. Three methodological differences explain the gap between Kaplan’s N* ∝ C^0.73 and Chinchilla’s near-C^0.5.

The learning-rate schedule. Kaplan’s runs used a cosine decay set to a fixed horizon, then read off intermediate checkpoints as if each were a completed run. A checkpoint halfway through a cosine schedule has not annealed, and its loss is systematically worse than a run that was scheduled to end there. That penalizes long-training runs and makes data look less valuable than it is. Chinchilla matched the decay length to each run’s own token budget.

Embedding parameters. Kaplan excluded embedding and unembedding parameters from N. For small models those are a large fraction of the total, so the smallest models were credited with far fewer parameters than they had, tilting the fitted slope. Chinchilla counted all parameters. Range compounded both: a narrower ladder of model sizes leaves less leverage to separate the two exponents.

The constrained optimization

Now the central derivation. Training compute for a dense transformer is well approximated by C ≈ 6ND — roughly 2 FLOPs per parameter per token forward and 4 backward. Fix a budget C and ask which (N, D) on that hyperbola minimizes loss. Form the Lagrangian J = L(N, D) + λ(6ND − C) and set both partials to zero:

∂J/∂N = −αA·N^−(α+1) + 6λD = 0
∂J/∂D = −βB·D^−(β+1) + 6λN = 0

multiply the first by N, the second by D, and equate 6λND:

       α · (A / N^α)  =  β · (B / D^β)

That last line is the whole optimum in one equation, and it is elegant: at the compute-optimal point the two reducible loss terms are held in a fixed ratio, weighted by their own exponents. Substituting D = C/(6N) and solving for N gives the closed form.