Fine-tuning is a different scaling regime
Pretraining scaling laws — Kaplan, Chinchilla — describe a model learning a distribution from nothing: loss falls as a power law in parameters N, data D, and compute C. Fine-tuning inverts the setup. The model already sits near a good general solution, and you are nudging it toward a narrower target with a comparatively tiny dataset — often 10^3 to 10^5 examples rather than the 10^{12} tokens of pretraining.
That changes what ‘more’ means. In pretraining, doubling data reliably lowers loss along a known slope. In fine-tuning, the base model has already absorbed most of the relevant structure, so the first hundred examples can move the metric more than the next ten thousand: the curve is far steeper at the start and flattens sooner. The right mental model is not ‘how much can this dataset teach?’ but ‘how much of what this task needs did pretraining already supply, and how little new data surfaces it?’
Effective data transferred: pretraining priced in data
Hernandez, Kaplan, Henighan and McCandlish (2021), Scaling Laws for Transfer, gave the sharpest way to quantify this. Their idea: measure what pretraining is worth by asking how much extra fine-tune data a from-scratch model would have needed to reach the same loss. They call that the effective data transferred, D_T.
Concretely, a pretrained model fine-tuned on D_F task examples performs as if it had been trained from scratch on D_F + D_T examples. The pretraining does not add task data literally — it adds effective task data. When D_F is small, D_T can dwarf it: pretraining is doing almost all the work and your handful of examples merely points the model at the right slice of what it already knows. This one reframing — pretraining denominated in fine-tune data — is what makes every later result about diminishing returns and model size fall into place.
The transfer power law and its exponents
Hernandez et al. fit the effective data transferred as a power law in both the fine-tune dataset size and the model size:
D_T = k * (D_F)^α * (N)^β
D_T = effective data transferred (extra data pretraining is worth)
D_F = size of the fine-tuning dataset
N = model size (non-embedding parameters)
α ≈ 0.38 (exponent on fine-tune data)
β ≈ 0.63 (exponent on model size)Two facts do the heavy lifting. First, β > 0: transfer grows with model size, so a larger base model extracts more effective data from the same task set. Second, α < 1: transfer grows sublinearly in fine-tune data. Double D_F and D_T rises by only 2^0.38 ≈ 1.30×. Because the transferred term grows slower than the data you add, pretraining’s share of the total shrinks as your dataset grows — the mathematical seed of diminishing returns.
A worked example
Suppose fine-tuning a model on D_F = 1,000 examples behaves like training from scratch on D_T = 100,000 effective examples. Pretraining is worth 100× your data — an enormous leverage, typical of the low-data regime.
Now grow the fine-tune set to D_F = 100,000, a 100× increase. With α ≈ 0.38, D_T scales by 100^0.38 ≈ 6.6×, reaching roughly 660,000. Absolute transfer rose, but relative transfer collapsed: pretraining went from being worth 100× your data to under 7× it. The effective total is now dominated by real task data, not by the head start. The lesson is quantitative: pretraining is a multiplier that fades as you accumulate task data, and it matters most in the regime small teams live in — hundreds to low thousands of labels.
Why bigger base models need less task data
The N^0.63 term has a direct operational reading. Take two pretrained models, one 10× larger. On the same fine-tuning set the larger model’s effective data transferred is higher by 10^0.63 ≈ 4.3×. It reaches a target loss with a small fraction of the examples the smaller model needs.
This is why practitioners repeatedly find that starting from a stronger, larger base is the highest-leverage decision in a data-scarce setting — it can beat collecting more labels. A bigger model has richer, more general features already in place, so a task is more likely to be a short interpolation away from something it can do rather than something it must learn from scratch. The corollary for budgets: when labels are expensive and a one-off fine-tune is cheap, spending on a better base model often dominates spending on annotation — a trade transfer scaling makes explicit rather than a matter of taste.
Diminishing returns and the high-data regime
Push D_F high enough and the story flips. Once your task data is large relative to the effective transfer, the pretrained model and a from-scratch model converge — both are mostly learning from your data, and the head start becomes a rounding error. Hernandez et al. describe this as the boundary between a low-data regime, where transfer dominates and pretraining is decisive, and a high-data regime, where it barely registers.
There is a sharper failure at the extreme called ossification: a small model pretrained on a very different distribution can fine-tune worse than the same model trained from scratch, because its weights are stuck in a basin that suits the pretraining task and resists being reshaped by abundant new data. Pretraining is not free insurance — it is a prior, and a strong prior only helps while you lack the data to overrule it.
Data quality dominates quantity
Because fine-tuning lives in the low-data regime where each example is doing outsized work, the composition of the set matters more than its raw size. The LIMA result — strong instruction-following from roughly a thousand carefully curated examples — is the canonical demonstration: past a modest point, adding mediocre examples moved the needle less than replacing them with cleaner ones. A noisy or inconsistent label is not neutral; it actively teaches the model the wrong mapping in a regime where the model is highly sensitive to each signal.
This refines the power laws rather than contradicting them: the transfer law assumes examples of a given quality, and degrading quality shifts the effective-data curve downward. For a fixed labeling budget, the scaling-aware move is fewer, higher-quality, more diverse examples covering the task’s hard cases, rather than bulk data that restates what the base model already handles and dilutes the gradient with noise.
Full fine-tuning versus PEFT scaling
Full fine-tuning updates every weight; parameter-efficient fine-tuning (PEFT) — LoRA, adapters, prefix tuning — freezes the base and trains a small add-on. They scale differently, and the crossover is governed by how much you need to move the model. When the target task is close to pretraining and the dataset is small — the common fine-tuning case — PEFT typically matches full fine-tuning, because the required change is genuinely low-rank and a small module has enough capacity to express it.
As you scale fine-tune data or push toward a task far from pretraining, the picture separates. Full fine-tuning keeps improving with data because it has the capacity to absorb it; PEFT can saturate once the change the task demands exceeds what the add-on’s limited parameters can represent. Recent studies of LoRA versus full fine-tuning report exactly this: comparable results on modest adaptation, a widening gap on large-data or hard-domain regimes where full fine-tuning’s extra degrees of freedom pay off.