The paper everyone cited and the fit almost no one re-ran

Hoffmann et al. estimated the compute-optimal split three independent ways, and the agreement between them is a big part of why the result was believed. Two of the three are essentially model-free: Approach 1 reads the loss-minimizing model size straight off training curves at each compute budget, and Approach 2 fits IsoFLOP profiles — loss versus model size along fixed-compute slices — and locates each valley. Both pointed at roughly equal scaling of parameters and data.

Approach 3 is different in kind. It fits a single closed-form law for loss as a function of parameters N and tokens D, then minimizes it analytically. It is the most elegant of the three and the one everyone quotes the constants from — and, it turns out, the one nobody had reproduced. The Epoch team’s contribution was simply to try, which is harder than it sounds when the underlying training runs were never released.

Advertisement

What Approach 3 actually estimates

The parametric law is the familiar three-term fit — an irreducible floor plus a finite-model penalty plus a finite-data penalty:

L(N, D) = E + A / N^α + B / D^β

The full derivation of how minimizing this under C ≈ 6ND yields N_opt ∝ C^a and D_opt ∝ C^b lives in the companion ‘Chinchilla Scaling’ article; here only the fitting matters. Hoffmann et al. estimated the five constants by minimizing a Huber loss between predicted and observed log-loss over their runs, using the L-BFGS-B optimizer from a grid of initial guesses. Their reported solution was E ≈ 1.69, A ≈ 406.4, B ≈ 410.7, α ≈ 0.339, β ≈ 0.285, giving optimal exponents a ≈ 0.454 and b ≈ 0.542. Those five numbers are the ones the whole community copied.

Advertisement

No data release, so reconstruct the figure

The first obstacle to replication was that Hoffmann et al. never published the loss values behind Approach 3. What they did publish was a scatter plot — Figure 4, the final-loss cloud used for the parametric fit. So the Epoch team did the only thing available: they digitized the figure, extracting the coordinates of every plotted point (roughly 240 runs across ten model sizes) by reading the underlying vector graphics and recovering each point’s (N, D, loss) triple from its position and color.

This is inherently lossy — you inherit the plot’s rounding and any overlap between markers — and the authors are careful to say so. But it is enough to ask a sharp question that the original paper never invited: if you take Hoffmann et al.’s own reported constants and draw the curve they imply, does it actually pass through the cloud of points in their own figure? The answer was no.

Finding one: the published fit misses its own data

When Epoch plugged the paper’s reported (E, A, B, α, β) back into the loss law and compared it against the reconstructed points, the fit was visibly poor — far worse than a five-parameter model fit to a few hundred points should be. That is a red flag on its own, but the more damning symptom was internal disagreement. Hoffmann et al.’s Approach 3 constants imply a compute-optimal ratio of roughly 70 tokens per parameter at Chinchilla scale, not 20.

That contradicts the paper’s own Approaches 1 and 2, and it contradicts how the Chinchilla model was actually trained (70B parameters on 1.4T tokens, a 20:1 ratio). In other words, the most-quoted set of constants in the paper was the outlier of the three methods, quietly inconsistent with the very model the paper is named after. The agreement everyone remembered was really an agreement between Approaches 1 and 2, with Approach 3 tagging along under a rounded headline.

Finding two: confidence intervals too tight to be real

The second problem was statistical. Hoffmann et al. reported the optimal exponents with startlingly narrow confidence intervals — on the order of a = 0.454 to 0.455, a width near 0.001. Epoch pointed out that intervals that tight are simply not compatible with the size of the experiment. To pin an exponent to a thousandth with the observed scatter, you would need something like 240 × 2116 ≈ 600,000 training runs; Hoffmann et al. describe having on the order of 400 to 500.

Part of this is mundane — reporting the exponents rounded to two decimal places manufactures a spuriously precise interval — but part is a genuine statistical point: the uncertainty on these exponents is real and much larger than advertised. A confidence interval three orders of magnitude too narrow is not a rounding footnote; it is the difference between ‘we have measured this’ and ‘we have estimated this, loosely.’

The bug: averaging the Huber loss instead of summing it

Why did the fit land in the wrong place? Epoch traced it to the optimization itself. Reconstructing the fitting procedure, they found that the original code appears to have averaged the per-point Huber losses rather than summing them. Averaging shrinks the objective by a factor of the number of data points, which drives the gradient magnitude and the loss scale far below what L-BFGS-B expects.

L-BFGS-B decides it has converged when the objective stops changing by more than a tolerance. With the loss scaled down by a couple of hundred, that tolerance is met almost immediately, so the optimizer terminates early — parked near its initialization rather than at the true minimum. The reported constants are essentially a half-finished optimization. It is a small, humble bug, the kind that lives in a hundred research repositories, and it is a vivid reminder that a published number is only as good as the loop that produced it.

The refit reconciles all three approaches

When Epoch re-ran the same parametric fit correctly — summed loss, a converging optimizer, the same functional form — the constants moved:

ConstantHoffmann (as reported)Epoch refit
E (irreducible)1.691.82
A406.4482.0
B410.72085.4
α0.3390.348
β0.2850.366
a (param exponent)0.4540.513 ± 0.02

The single biggest change is B, which jumps roughly five-fold, and β, which rises to meet α. Because a = β / (α+β), the near-equal exponents push a to about 0.51 — genuinely balanced scaling — and the implied ratio back to roughly 20 tokens per parameter. Approach 3, once fixed, stops being the outlier and agrees with Approaches 1 and 2 and with how Chinchilla was trained.

Why the parametric fit is intrinsically fragile

The bug is fixable, but it exposed something deeper: this fit is ill-conditioned. The two power-law terms A/N^α and B/D^β can partly impersonate each other over the observed range, so a big change in B can be compensated by a small change in β with almost no change in the fitted loss. The objective has a long, shallow valley rather than a sharp bowl.

In that regime the answer you get depends on things that should not matter: the Huber transition point δ, the grid of initial seeds, the stopping tolerance, even the loss scaling that caused this bug. Two honest researchers running defensible procedures can land on exponents that differ in the second decimal — the range that decides ‘20 tokens per parameter’ versus ‘40.’ The lesson is not that the law is wrong but that its coefficients deserve error bars and sensitivity checks, not the false precision of a single quoted decimal.