What an inference-time scaling law measures

A pretraining scaling law fixes the recipe and varies the model: loss L(N, D) as a function of parameters and data. An inference-time scaling law does the opposite — it fixes the model and varies how much compute you spend answering a single query. The x-axis is test-time compute C (measured in generated tokens, in number of samples N, or in FLOPs per query); the y-axis is a task metric such as accuracy or pass rate.

The reason this is interesting at all is that a frozen model is not a single function from prompt to answer. Sampling is stochastic, so the model defines a distribution over answers, and different reasoning paths reach different conclusions. Spending more compute lets you draw more from that distribution and keep the best draw, or follow a single reasoning path further before committing. Both turn raw compute into accuracy without touching the weights — and both obey surprisingly clean laws, which is what makes test-time compute a budget you can plan with rather than a black box.

Advertisement

The knobs: what inference compute actually buys

‘More inference compute’ is not one lever but a family of them, and they scale differently. The main axes:

Sequential depth — a longer chain of thought. The model emits more intermediate tokens before its final answer, so compute grows with the length of a single trajectory. This is what ‘reasoning’ models spend on.

Parallel width — draw N independent samples for the same prompt, then pick one. Compute grows linearly in N. This is best-of-N, self-consistency, and majority voting.

Search — expand a tree of partial solutions, scoring and pruning branches (beam search, MCTS, lookahead). This mixes depth and width under the control of a value estimate.

All three ultimately cash out in generated tokens, so you can put them on a common compute axis. But they have different exchange rates between compute and accuracy, and — crucially — the best mix depends on the problem, a point the compute-optimal section makes precise.

Advertisement

Best-of-N and the coverage curve

The cleanest case to reason about is best-of-N: sample N candidate answers independently, then select one. Split the metric into two questions. First, coverage: does at least one of the N samples contain a correct answer? Second, selection: can you actually pick the correct one? Coverage is the ceiling; selection is how close you get to it.

Coverage is what the literature calls pass@N under oracle selection. If each sample is correct independently with probability p, then the probability that all N fail is (1 - p)^N, so

pass@N = 1 − (1 − p)^N

This rises fast and saturates toward 1: the error (1 - p)^N decays geometrically in N. If a perfect verifier could always spot the correct sample, this would be your accuracy. Real selection — majority vote, a learned verifier, a reward model — falls short of the oracle, so measured best-of-N accuracy tracks below the coverage curve and eventually flattens once selection, not coverage, becomes the bottleneck.

Why the empirical curve is log-linear

The geometric formula predicts error falling linearly in N, which would look explosive. In practice, plotting accuracy against N on a log x-axis gives a much tamer, near-straight line over a wide range:

Accuracy(C) ≈ a + b · log(C)

Each doubling of compute adds a roughly constant b · log(2) to accuracy. Why the gap between the geometric ideal and the log-linear reality? Because a single per-sample probability p is a fiction. A real benchmark is a mixture of problems with wildly different difficulties — some with high p, some near zero. Easy problems are solved after a handful of samples and stop contributing; the marginal gains come from progressively harder slices of the distribution. When the density of problem difficulty is spread roughly evenly across orders of magnitude of ‘samples needed’, each multiplicative increase in N unlocks about the same increment of the benchmark — which is exactly a log-linear law. The clean per-problem geometry averages into a clean per-benchmark logarithm.

A worked numeric example

Take a problem the model solves with per-sample probability p = 0.4. Under oracle selection, coverage climbs like this:

N (samples)Error (1−p)^Npass@N
10.6000.400
40.1300.870
160.00280.997
64~6×10^-8≈1.000

Coverage is nearly perfect by N = 16. But this is the ceiling, not the achieved score. Suppose your selector — say majority vote — only picks the correct answer 70% of the time when one exists in the pool. Then realized accuracy at N = 16 is closer to 0.997 × 0.70 ≈ 0.70, and pushing N higher barely helps because you are no longer coverage-limited — you are selection-limited. This is the single most important reason best-of-N curves bend: the compute keeps finding correct answers you cannot reliably identify.

Compute-optimal test-time allocation

Given a fixed per-query compute budget C, how should you spend it? The question is real because the knobs compete: tokens spent going deeper on one chain are tokens not spent going wider across samples. Snell et al. (2024) framed this as choosing, per query, the allocation that maximizes expected accuracy at budget C — the compute-optimal strategy.

Their empirical finding is that the best allocation is difficulty-dependent. On easy and medium problems, spending budget on sequential refinement — revising a single answer — beats brute-force parallel sampling; the model is close and needs a nudge, not a lottery. On the hardest problems, broad parallel search wins, because the model needs many independent shots to cover an answer at all. Adapting the strategy to a predicted difficulty can match a fixed strategy’s accuracy with several times less compute. The practical lesson: do not pick one test-time method globally — route compute by how hard each query looks.

The train-versus-inference compute tradeoff

The deepest consequence of inference scaling is that training compute and inference compute are partly interchangeable. A weaker model that thinks longer can reach the accuracy of a stronger model that answers in one shot. Snell et al. showed that on many problems a smaller model under compute-optimal test-time scaling matches or beats a model roughly 14× larger evaluated greedily — provided the questions are not in the hardest tier, where extra pretraining still wins.

This mirrors a result Andy Jones (2021) found in board games: for AlphaZero-style agents, train compute and test-time search compute trade off along a smooth, roughly log-linear frontier — you can buy a fixed increment of strength with either. The existence of such a frontier is what makes ‘let it think’ a genuine engineering axis rather than a gimmick: you are moving along a known exchange rate between two kinds of compute, and you get to pick the cheaper one for your deployment.