Where the clean power law stops describing reality
A power law is a strong claim: the same multiplicative rule applies everywhere, so a fit measured at 10^19 FLOPs tells you what happens at 10^23. Every failure mode here violates that assumption somewhere different.
They sort into four kinds. Input violations — the D you feed the law is not what the law means, because your tokens are repeats, or because N is ambiguous in a sparse model. Regime violations — the exponent changes partway along the curve, so one straight line in log-log space is the wrong shape. Output violations — you fit loss but you ship accuracy, and the map between them is nonlinear. And methodological violations — the fit silently encodes choices that will not hold in the run you are trying to predict.
Data-constrained scaling: a repeated token is not a token
The law treats D as tokens seen, but implicitly means unique tokens: fresh samples each carrying new information. Once you exhaust your corpus and start a second epoch that assumption quietly breaks — the gradient from a token the model has already fit contributes far less than a novel one, and eventually nothing but memorization pressure.
The empirical finding is more encouraging than intuition suggests. Repeating a corpus for up to roughly four epochs costs remarkably little: the loss curve tracks the fresh-data curve closely enough that those repeated tokens are nearly free for planning purposes. Past that, the value of each additional epoch decays roughly exponentially, and by around 16 epochs extra passes buy essentially nothing — the curve flattens onto a floor set by how much information your unique data actually contains. This is why ‘how many tokens do we have’ became a harder constraint than ‘how many GPUs do we have.’
The effective-data model, worked
The clean way to keep using the ordinary law is to stop pretending repeats are fresh and instead compute an effective token count to substitute for D. With U unique tokens and R extra repetitions beyond the first pass, the fitted form used in the data-constrained scaling literature is:
D_eff = U + U · R* · (1 − exp(−R / R*))
R = extra repetitions (total epochs − 1)
R* ≈ 15 = fitted decay constant for repeated data
as R → ∞: D_eff → U · (1 + R*) ≈ 16·UTwo consequences. There is a hard ceiling: no amount of re-reading extracts more than about 16× your unique data’s worth. And a worked case — take U = 10B. Four epochs (R = 3) gives D_eff = 10 + 154·(1 − e^(-0.195)) ≈ 37B against 40B nominal, 93% efficiency. Sixteen epochs (R = 15) gives ≈ 106B against 160B nominal — 66%, with the marginal epoch worth only e^(-0.97) ≈ 38% of a fresh one. Treat the constant as a fitted parameter, not a law of nature; the shape is the durable part.
Mixture of experts: which N do you even plug in?
A sparse mixture-of-experts layer replaces one feed-forward block with many and routes each token to a small subset. The model now has three ‘sizes,’ and conflating them is how most MoE scaling arguments go wrong:
| Quantity | What it governs |
|---|---|
Total params N_total | Memory footprint, checkpoint size, what you must hold to serve |
Active params N_act | FLOPs per token — roughly 2·N_act forward |
| Effective params | Quality — the dense model it behaves like |
The dense law folds all three into one N because in a dense model they coincide. In a sparse one they diverge sharply. A rough back-of-envelope people quote — a heuristic, not a fitted scaling law — is that quality sits near the geometric mean, sqrt(N_total · N_act). For a Mixtral-style 8×7B with top-2 routing: 46.7B total, ~12.9B active, sqrt(46.7 · 12.9) ≈ 25B. It costs the FLOPs of a 13B model, occupies the memory of a 47B model, and performs like a 25B one.
Sparsity as a third axis
Properly, MoE adds a dimension to the law rather than breaking it. Fitted routed-model laws express loss in terms of active parameters and expert count, with a saturating gain in experts: 1 → 8 is a large win at fixed FLOPs, 8 → 64 a smaller one, and beyond that the curve bends toward the dense law at the same N_act. Sparsity buys a translation along the loss axis, not a change in the exponent.
Two corollaries. Comparing MoE to dense is only meaningful once you say which budget is fixed — equal FLOPs flatters the MoE, equal memory flatters the dense model, and equal quality is the only comparison a user feels. And granularity matters: many small experts with a wider top-k generally beat a few large ones at the same active count, because you get finer specialization per unit of compute. ‘How sparse’ is a design variable with its own optimum, not a free win.
Emergent abilities and the measurement critique
The most-discussed apparent break is emergence: a capability sits at chance across several orders of magnitude, then jumps sharply above some scale. If real, that is fatal for extrapolation — no smooth fit predicts a cliff.
The measurement critique argues the cliff is often in the metric, not the model. Exact-match scoring is a harsh nonlinear transform of per-token competence. If a task needs five tokens right and per-token accuracy is p, exact match is p^5. As p improves smoothly 0.3 → 0.5 → 0.7 → 0.9, exact match reads 0.002 → 0.03 → 0.17 → 0.59 — flat, flat, then a hockey stick, from a perfectly linear underlying trend. Re-score the same runs with a continuous metric (token edit distance, Brier score, log-likelihood of the correct answer) and many published emergence curves become smooth and predictable.
Be careful how far you take this. It explains the shape of many curves; it does not prove nothing changes qualitatively with scale, nor tell you which capability will cross a usability threshold — a real, discontinuous fact even when the underlying loss is smooth.