What Kaplan actually measured
The study fixed almost everything except scale. Decoder-only transformers were trained autoregressively on WebText2, and the reported quantity is the cross-entropy loss in nats per token on held-out data — not accuracy, not a downstream benchmark, just the language-modeling loss the network is directly optimizing.
A crucial bookkeeping choice: N counts non-embedding parameters. Embeddings are excluded because they scale with vocabulary and context, not with reasoning capacity, and including them muddies the trend at small sizes. With that convention, three quantities matter: N (model size), D (dataset size in tokens), and C (compute in FLOPs, well approximated by C ≈ 6ND — roughly two FLOPs per parameter forward and four backward). Everything in the paper relates the loss to one of these three, with the other two held non-limiting.
The three power laws
When one resource is the bottleneck and the others are abundant, the loss follows a clean power law. Kaplan reports all three:
L(N) = (N_c / N)^α_N α_N ≈ 0.076, N_c ≈ 8.8 × 10^13
L(D) = (D_c / D)^α_D α_D ≈ 0.095, D_c ≈ 5.4 × 10^13
L(C) = (C_c / C)^α_C α_C ≈ 0.050 (compute-optimal frontier)Read L(N) as: with enough data and training, a model of size N reaches this loss. L(D): a sufficiently large model, early-stopped, extracts this much from D tokens. L(C): if you spend compute C optimally, this is the best loss achievable. The exponents are small — a factor of ten more parameters cuts the loss by only 10^-0.076 ≈ 0.84, about 16 percent — which is exactly why progress demanded such enormous increases in scale.
How to read an exponent that small
A power law L = (x_c / x)^α is a straight line of slope -α when you plot log L against log x. The smallness of α is the whole economic story of large models: because α_N ≈ 0.076, each halving of the loss gap requires roughly 2^(1/0.076) ≈ 9,000× more parameters.
This is diminishing returns made quantitative, but it cuts both ways. The returns shrink, yet they never stop and they never surprise you: the line does not bend, so the next order of magnitude is as predictable as the last. That predictability is what made scaling an engineering decision rather than a gamble — you could read the expected loss off the fitted line and justify a seven-figure training run before launching it.
The joint law and the overfitting boundary
Real training is bottlenecked by N and D together, and Kaplan fits a single equation that unifies them:
L(N, D) = [ (N_c / N)^(α_N / α_D) + D_c / D ]^α_DThe structure is intuitive. When D is huge the second term vanishes and you recover L(N); when N is huge the first term vanishes and you recover L(D). The interesting content is the cross term: to keep the data penalty from dominating as you grow the model, you must scale D ∝ N^(α_N / α_D) ≈ N^0.74. Because that exponent is less than one, Kaplan concluded that data should grow sublinearly with model size — a bigger model needs more tokens, but proportionally fewer per parameter. That single sublinear exponent is the seed of the whole ‘bigger, not longer’ doctrine, and — as we will see — it is precisely the number the later correction moved.
Compute-optimal allocation: the N^0.73 rule
Fix a compute budget C ≈ 6ND. You may spend it on a large model trained on few tokens, or a small model trained on many. Which minimizes the loss? Kaplan differentiates the joint law under the constraint and finds that the optimal model size grows as a strong power of compute:
N_opt ∝ C^0.73 (params take almost all new compute)
D_opt ∝ C^0.27 (tokens grow slowly)
B_opt, steps grow slowly tooThe 0.73 exponent is the paper’s most consequential single number. It says that when your compute budget grows 10×, you should make the model about 10^0.73 ≈ 5.4× bigger but feed it only 10^0.27 ≈ 1.9× more data. Taken literally, it is an instruction to build gigantic models and under-train them — stop well before the data is exhausted, because another dollar buys more loss reduction as parameters than as tokens. The field followed this faithfully: GPT-3, at 175 billion parameters trained on only ~300 billion tokens, is almost a direct expression of the Kaplan allocation.
The critical batch size
Kaplan’s second major result governs how fast you can train, not just how well. Building on the gradient-noise-scale theory, they define a critical batch size B_crit that marks the transition between two regimes:
B_crit(L) ≈ B_* / L^(1/α_B) B_* ≈ 2 × 10^8 tokens, α_B ≈ 0.21Below B_crit, doubling the batch nearly halves the number of optimizer steps at almost no extra total compute — trading hardware parallelism for wall-clock time essentially for free. Above it, larger batches waste compute for diminishing speedups. Crucially, B_crit depends only on the loss, not on model size, so it grows as the loss falls: late in training, and for stronger models, you can profitably use much larger batches. Pick a batch near B_crit to sit at the efficient frontier between minimum steps and minimum total FLOPs.
Larger models are more sample-efficient
A counterintuitive corollary falls out of the same fits: bigger models learn more from each token. Plot loss against tokens processed and the large models sit below the small ones at every point, not just at convergence — for a fixed target loss, a larger model reaches it having seen fewer tokens.
This is why the compute-optimal recipe stops training early. If you are budget-constrained, the fastest path to a given loss is a large model trained briefly, not a small model trained to convergence. It also reframes overfitting: the sample-efficient large model hits its target before it has cycled through enough data to memorize, so the classic overfitting regime is simply never entered in the compute-optimal setting — and the joint law quantifies when it would begin, letting you size a dataset to a model rather than discovering the mismatch after a wasted run.
The infinite-data and infinite-compute limits
The two single-variable laws are best understood as opposite idealized corners of the training space. L(N) is the infinite-data limit: hold data and steps non-limiting and ask what a model of size N can ultimately represent. It is a statement about model capacity, a floor set by parameters alone.
L(C) is the compute-optimal frontier: at each budget it assumes you have allocated N, D, and batch size optimally, so it traces the best-possible loss per FLOP rather than any single run. The gap between the naive path (train one fixed model to convergence) and this frontier is exactly the efficiency the allocation rule recovers. Kaplan extrapolated the two limits toward each other and noticed they must eventually collide — the point where compute-optimal training would consume essentially all available text — flagging it as the boundary where these simple power laws must break down, and foreshadowing the data-wall conversations of years later.