The setup: generate many, keep one

Test-time scaling gives a model more than one shot. The generator samples N candidate solutions y_1 … y_N for a prompt (temperature > 0 so they differ), and a selector chooses one to return. Three selectors dominate: majority vote (self-consistency — return the most common final answer), best-of-N with a verifier (return the candidate the verifier scores highest), and weighted combinations. The verifier is a learned scorer V(x, y) → ℝ — often a reward model or a model fine-tuned to predict ‘is this solution correct?’. The question this article answers: how does accuracy grow with N, and what caps it?

Advertisement

Two ceilings: coverage and selection

Every best-of-N system is squeezed between two numbers. Coverage (also called pass@N) is the chance that at least one of the N samples is correct — the best you could possibly do with a perfect selector. If a single sample is correct with probability p, then for independent samples

coverage(N) = pass@N = 1 − (1 − p)^N

which rises fast toward 1. The selection accuracy is what you actually get: the chance the selector picks a correct one. A perfect verifier achieves coverage; a useless verifier (random pick) achieves p. Real verifiers sit in between, and the gap between the coverage curve and the achieved curve is the verifier’s tax.

Advertisement

The best-of-N accuracy curve

Model the verifier as assigning a score to every candidate, correct or not. Best-of-N returns the argmax score. Accuracy is the probability the top-scored candidate is correct, which depends on how well the verifier separates correct from incorrect solutions. With a strong verifier, accuracy tracks the coverage curve and keeps climbing with N. With a weak, noisy verifier, accuracy rises then falls: past some N, the growing pile of wrong candidates eventually contains one that the noisy verifier scores higher than any correct one (an adversarial-sampling effect — the same Goodhart failure as reward over-optimization). So more samples is not always better; the optimal N depends on verifier quality.

A worked example

Take p = 0.3 (30% single-shot). Coverage climbs as:

N12481632
coverage = 1−0.7^N0.300.510.760.940.9971.00

A perfect verifier would deliver that top row — 94% at N=8. Majority vote lands lower, near the model’s single most-likely answer, and saturates at the model’s inherent consistency (often well under coverage). A good but imperfect verifier might realize, say, 0.30 → 0.48 → 0.63 → 0.74 → 0.80 — beating majority vote by rescuing minority-correct answers, but short of the 0.94 ceiling. The three curves (coverage > verifier-BoN > majority vote) are the whole story.

Why a verifier beats majority vote

Majority vote can only be right if the correct answer is the plurality answer. On hard problems the model is often confidently wrong in a consistent way — the same seductive mistake recurs, and it wins the vote. A verifier is not bound by popularity: it can select a correct solution that appeared only once, as long as it scores it above the wrong ones. This is why verification is generically easier and more powerful than generation for problems where checking is easier than solving (math, code with tests, formal proofs). It is also why majority vote plateaus while a strong verifier keeps benefiting from more samples.

Outcome vs process reward models

Verifiers come in two flavors. An outcome reward model (ORM) scores only the final answer — one label per solution. A process reward model (PRM) scores each reasoning step, giving dense feedback along the chain. A PRM is more expensive to train (needs step-level labels) but is stronger: it catches a solution that reaches the right answer by faulty reasoning, and it can prune bad branches early during search rather than only judging completed solutions. Lightman et al.’s ‘Let’s Verify Step by Step’ found PRMs substantially outperform ORMs at best-of-N on hard math. The PRM step scores are typically aggregated (product or minimum of per-step correctness probabilities) into a single solution score for selection.

Verifier-guided search

Best-of-N verifies completed solutions. A PRM lets you verify as you go, turning selection into search: expand a tree of partial solutions and use the step scores to decide which branches to keep. Beam search over reasoning keeps the top-b partial chains by PRM score at each step. Tree search / MCTS goes further, balancing exploration and exploitation. Guided search spends compute where it matters (promising branches) instead of sampling N full solutions blindly, so it often reaches a target accuracy with fewer generated tokens — the same idea behind reasoning models that ‘think’ along a searched trajectory.

The compute split: samples vs verification

Inference budget is finite, so you must divide it. Generating N samples of length T costs about N·2P·T FLOPs for a P-parameter generator; scoring them with a verifier of size P_v adds roughly N·2P_v·T. A cheap verifier (small ORM) barely adds cost, so N can be large; a full-size PRM run at every step can cost as much as generation itself. The practical questions: is your verifier good enough that more N still helps (or does accuracy roll over)? And at a fixed budget, is it better to draw more samples, or to spend that compute making the verifier bigger/better? Empirically, verifier quality usually buys more than raw sample count once N is moderate.