Two knobs, one budget
Training a language model spends compute in one place that matters for scaling: pushing tokens through a network and updating its weights. The total is governed by two numbers — the parameter count N and the training-token count D. A fixed budget C ties them together, so you cannot maximize both. Spend it all on a huge model and you can afford only a few passes of data, leaving it under-trained; spend it all on data and the model is too small to use that data. Neither extreme is efficient.
The compute-optimal question is the constrained optimum between those failures: for a given C, which pair (N, D) reaches the lowest loss? Answering it needs two ingredients — a formula relating C to N and D (the constraint), and a model of how loss depends on them (the objective). With both, the optimum falls out of one clean piece of calculus — starting with the constraint.
The 6ND compute identity
Where does C ≈ 6ND come from? Start with one matrix multiply. Feeding a token’s activation through a weight matrix with P entries costs one multiply and one add per entry — 2 FLOPs per parameter. Summed over every weight matrix, the forward pass costs about 2N FLOPs per token.
The backward pass does about twice that work. For each layer it computes two gradients: one with respect to the inputs (to keep propagating the error) and one with respect to the weights (to update them). Each is a matmul the size of the forward one, so backprop costs about 4N FLOPs per token. Add them:
forward ≈ 2N FLOPs / token
backward ≈ 4N FLOPs / token
-----------------------------------
per token ≈ 6N FLOPs
over D tokens: C ≈ 6 · N · DThat is the whole identity: 6 FLOPs per parameter per token, times N parameters, times D tokens.
What the identity quietly ignores
C ≈ 6ND is an approximation worth knowing the blind spots of. It counts only the dense matrix multiplies in the linear layers — the attention projections and feed-forward blocks — which genuinely dominate a transformer’s FLOPs. It omits the O(N_seq^2) attention-score computation, embedding lookups, layer norms, activations, and the softmax; at typical sizes those are lower-order corrections, usually a sub-10–20% adjustment. The one to watch is the quadratic attention term, which stops being negligible at very long context. There is also a subtlety in N: the identity is cleanest when it counts non-embedding parameters, the ones doing per-token matmul work. Even so, it is accurate enough that everyone uses it — it turns a messy hardware question into two numbers.
Modeling the loss: a power law in N and D
The objective needs a functional form for loss in terms of N and D. Empirically, the loss of well-trained transformers follows a strikingly regular power law, captured by the parametric form used across the scaling-law literature:
L(N, D) = E + A / N^α + B / D^βRead it term by term. E is the irreducible loss — the entropy of language itself, a floor no model or dataset can cross. A / N^α is the finite-model penalty, the price of too few parameters, shrinking as N grows. B / D^β is the finite-data penalty, the price of too few tokens, shrinking as D grows. The exponents α and β are the curvatures — how quickly each penalty melts away as you invest. They will decide the optimal split, so hold onto them.
Setting up the Lagrangian
The problem is now fully specified: minimize L(N, D) subject to 6ND = C. This is a textbook constrained optimization, and the tool is a Lagrange multiplier λ. Fold the constraint into the objective and take derivatives:
ℒ(N, D, λ) = E + A/N^α + B/D^β + λ(6ND − C)
∂ℒ/∂N = −αA / N^(α+1) + 6λD = 0
∂ℒ/∂D = −βB / D^(β+1) + 6λN = 0
∂ℒ/∂λ = 6ND − C = 0The two stationarity conditions are intuitive once rearranged: at the optimum, the marginal loss reduction per extra FLOP is equal whether you spend that FLOP on parameters or on tokens. If one knob bought more loss-per-FLOP, you would shift budget toward it — so at the balance point they are equal, and that shared value is λ.
Solving it: the balance condition
Eliminate λ by solving each first-order condition for 6λ and setting the two expressions equal:
6λ = αA / (N^(α+1) · D) = βB / (D^(β+1) · N)
cross-multiply, cancel one N and one D:
αA · D^β = βB · N^α (balance condition)This is the heart of the result: at the optimum the two shrinking penalties are locked in a fixed ratio, α·(A/N^α) = β·(B/D^β) — you never fully starve either. The balance condition alone fixes the shape, forcing N ∝ D^(β/α), but not the scale. To pin down actual sizes, combine it with 6ND = C: two equations, two unknowns.
The general result: N and D as powers of C
Substituting N ∝ D^(β/α) into ND = C/6 and solving gives the compute-optimal allocation — each a power of the budget:
N* = G · C^a , a = β / (α + β)
D* = H · C^b , b = α / (α + β)
a + b = 1 (forced by ND ∝ C)Three takeaways. First, the exponents come entirely from the loss curvatures α and β — the geometry of the surface, not the constants A, B, E (which only set the prefactors G, H). Second, note the cross-over: a steeper data-curvature β pushes budget toward N, and vice versa. Third, a + b = 1 is no coincidence but a consequence of ND ∝ C: double the compute and both N* and D* grow so their product doubles.
The geometry: iso-loss and iso-FLOP contours
The calculus has a clean picture in (log D, log N) space. A fixed budget 6ND = C becomes log N + log D = const — a straight line of slope −1, an iso-FLOP line. A fixed loss L(N, D) = const traces a convex iso-loss contour, bowing away from the origin because you can trade some N for some D at the same loss.
For a given budget you want the lowest loss reachable on that line — the point where the iso-FLOP line just touches the best contour it can, the point of tangency. That tangency is exactly what the Lagrangian computes: λ is the shared slope where the two families kiss. Sweep the budget across many parallel iso-FLOP lines, collect the tangency points, and they trace a straight line in log-log space — the compute-optimal frontier, with slope b/a = α/β.