The curve everyone quotes hides a variable

The compute-optimal scaling law says test loss falls as a power of the compute you spend: roughly L(C) ≈ E + (C_c / C)^α_C, with a matching form in model size N and token count D, L(N, D) = E + A/N^α + B/D^β. The Chinchilla fit puts α ≈ 0.34 and β ≈ 0.28, and the E term is the irreducible entropy of the data. Every one of those constants — A, B, C_c — is measured on a particular corpus. Change the data distribution and you re-fit the constants.

That is the whole point. Adding more tokens moves you along a fixed curve toward its floor. Improving the data redraws the curve: the B/D^β term shrinks for every token count, so the same compute buys a lower loss. Quantity walks down the slope you were handed; quality hands you a better slope. Treating data quality as a fixed background assumption — the mistake baked into a bare reading of the law — hides the single highest-leverage lever a small team actually controls.

Advertisement

Quality as an effective-data multiplier

The cleanest way to reason about curation is to fold it into a single number. Say a curated corpus is worth q times a baseline web corpus token-for-token, so the model behaves as if it trained on D_eff = q · D tokens. The data-limited loss term becomes B / (q · D)^β. To hit a target loss you now need only D = D_eff / q real tokens, and since training compute is C ≈ 6 · N · D at fixed model size N, your compute bill falls by roughly the same factor q.

target data-term:  B / D_eff^β         (D_eff = effective tokens)
quality q:         D_eff = q · D           (q > 1 for better-than-web data)
real tokens:       D    = D_eff / q
compute (fixed N): C    ≈ 6 · N · D  →  drops ~q×

worked: need D_eff = 300B effective tokens for the target loss.
  baseline web,  q = 1.0  →  D = 300B raw tokens
  curated data,  q = 2.5  →  D = 120B raw tokens  (2.5× less compute)

The multiplier is an approximation — q is not constant across the whole curve and it saturates — but it captures the mechanism exactly: better data is indistinguishable, to the loss curve, from simply having more of the baseline data.

Advertisement

Textbooks are all you need: the Phi evidence

The sharpest demonstration is Microsoft’s Phi line. phi-1 (1.3B parameters) was trained on only about 7B tokens of ‘textbook-quality’ data — a filtered slice of web code plus model-generated textbook and exercise text — and reached roughly 50% pass@1 on HumanEval, rivaling models trained on one to two orders of magnitude more tokens. phi-1.5, phi-2, and phi-3 pushed the same thesis into general reasoning: a small model on curated data punches far above its parameter and token class.

In effective-data-multiplier terms, the Phi recipe is a large q. The authors argue that most web text is a poor teacher — boilerplate, noise, and low-density prose — while text that is self-contained, correct, and pedagogically ordered delivers far more signal per token. The provocative title ‘textbooks are all you need’ overstates it (you still need scale and breadth), but the empirical point stands: curation and targeted synthetic generation moved the curve enough that a sub-2B model did what much larger models did, at a fraction of the training compute.

Deduplication: the cheapest multiplier

Before any clever filtering, the highest return-on-effort quality win is removing duplicates. Large web scrapes contain enormous redundancy — boilerplate, mirrored pages, near-identical documents. Lee et al. (2021) showed that deduplicating training data makes language models better: models trained on de-duplicated data reach the same or better accuracy in fewer steps, emit far less memorized training text verbatim, and need fewer parameters to match a baseline.

The mechanism fits the multiplier picture. A document repeated 40 times does not give you 40 tokens’ worth of signal; it gives you one document’s signal plus 39 wasted passes that push the model toward memorization instead of generalization. Deduplication — exact-match plus near-duplicate detection via MinHash / LSH on n-gram shingles — raises the unique-information density of the corpus, so every token the optimizer sees carries more novelty. It costs a one-time preprocessing pass and no training compute, yet it can meaningfully lift q. It is the first thing any serious data pipeline does, and the easiest to under-rate.

Filtering and curation: FineWeb-Edu and DCLM

Beyond dedup, the modern lever is classifier-based filtering: train a lightweight model to score how ‘valuable’ a document is, then keep the high-scoring tail. FineWeb-Edu applies an educational-quality classifier to Common Crawl, keeping text that looks like it teaches something. Models trained on it reach a given accuracy on knowledge and reasoning benchmarks (MMLU, ARC) with roughly an order of magnitude fewer tokens than raw Common Crawl — a large effective-data multiplier, bought purely by selection.

DCLM (DataComp-LM) turned this into a controlled competition: fix the model and compute, vary only the data pipeline, and measure. Its best filtered subset, DCLM-Baseline, trained competitive models with several times less compute than weaker-filtered baselines, confirming that the data-processing recipe — not just token count — is a first-class scaling variable. The takeaway across both: a good quality classifier plus aggressive deduplication is, in effect, a compute multiplier you apply once, offline, before a single gradient step. (Exact multipliers vary by benchmark and setup; treat the figures as approximate.)