Two knobs: chain length and sample count

Test-time compute for a reasoning model comes in two flavors. The first is depth per chain — how many tokens of scratch work the model writes before its answer — which is serial computation: each step conditions on the last, so the model can carry state, check a case, and backtrack. The second is breadth — how many independent chains you sample — which is parallel compute: K generations that never see each other, later collapsed into one answer.

The two hit different walls. Depth is bounded by the model losing coherence over very long outputs; breadth is bounded by how often the model can reach the correct answer at all. This article treats both, but the quantitative heart is breadth via self-consistency, where the accuracy-vs-compute curve has a clean, computable shape.

Advertisement

Self-consistency: sample many, keep the majority

Greedy decoding gives one deterministic chain; if it takes a wrong turn early, the answer is wrong with no recovery. Self-consistency replaces that with a vote: sample K chains at nonzero temperature so they diverge, extract the final answer from each, and return the one that appears most often. The intuition is that many correct reasoning paths reach the same right answer, while wrong paths scatter across many different wrong answers — so the correct answer accumulates a cluster of votes while errors split theirs and cancel out.

Crucially, self-consistency marginalizes over the reasoning and votes only on the final answer. It needs no verifier, no reward model, no extra training — just a way to parse the answer and count. That separates it from tm_verifier_scaling, where a learned scorer weights or selects chains; here every chain gets one equal vote. It is the cheapest possible aggregation and works surprisingly well — but its ceiling is set entirely by the raw sampling distribution, which the next sections make precise.

Advertisement

The vote, formally

Fix a problem. Let a single sampled chain produce final answer value v with probability p_v — the model’s answer distribution under the sampling temperature, with Σ_v p_v = 1. Draw K chains i.i.d., let n_v be how many landed on value v, and return the mode:

a_1, ..., a_K  ~ i.i.d.  answer distribution {p_v}
n_v  = Σ_{i=1..K}  1[a_i = v]          (vote count for v)

answer_hat = argmax_v  n_v                  (majority / plurality vote)

solved(K)  = 1[ argmax_v n_v == v* ]        (v* = correct value)

By the law of large numbers, as K → ∞ the fractions n_v / K → p_v, so the vote converges to argmax_v p_v — whichever answer the model is most likely to emit. So self-consistency eventually solves a problem if and only if the correct answer is the modal answer (p_{v*} > p_v for all wrong v). No amount of sampling fixes a problem where a wrong answer is more probable than the right one. That single fact is the whole story of the curve: sampling denoises, it does not create knowledge.

The accuracy-vs-K curve: diminishing returns to a plateau

Now sweep K. For one problem where the correct answer is modal, the probability the vote is correct rises with K but with shrinking marginal gains: the first few extra samples move the needle a lot, then each doubling of K adds less. Averaged over a dataset, Acc_SC(K) climbs steeply at small K, bends over, and flattens. It approaches a ceiling equal to the fraction of problems where the correct answer is the model’s modal answer — call it the coverage ceiling. Problems where the model is confidently wrong (a wrong answer is modal) sit permanently below the line; extra votes only entrench the wrong plurality.

This is why practitioners quote self-consistency at modest budgets — often K around 5 to 40 — and rarely beyond. The curve is concave: you capture most of the gain early, and past the knee you pay linearly more compute (each vote is a full generation) for asymptotically nothing. Two models can share the same greedy accuracy yet plateau at very different ceilings, because the ceiling depends on the shape of the answer distribution, not just its top value. Raising it is a training problem, not an inference one.

A worked accuracy-vs-K example

Take a problem where each sampled chain is correct with probability p = 0.6, and the remaining 0.4 piles onto a single tempting wrong answer. Correct is modal, so the vote should win as K grows. It is correct when a majority of the K samples are correct, i.e. P(Binomial(K, 0.6) > K/2):

K     P(vote correct)      gain vs previous
 1        0.600                 —
 3        0.648               +0.048
 5        0.683               +0.035
 9        0.733               +0.050
21        0.826               +0.093  (over 12 extra samples)
51        0.923
∞     1.000   (correct is modal, so the vote converges to it)

Two things stand out. The payoff is real but slow — K=1 to K=5 buys +8 points, while the next 16 samples (5→21) buy +14 for far more compute: textbook diminishing returns. And this problem is a lucky one. Rerun it with p = 0.45 against a 0.55 wrong answer and the same math runs in reverse (P(vote correct) → 0) — that problem drags the dataset ceiling down no matter how large K gets, the plateau in one line of algebra.

Longer CoT helps — until it hurts

Depth has its own curve, and it is not monotonic. A transformer does fixed computation per token, so the only way it performs a calculation deeper than one forward pass is to spread it across more tokens — the written chain is external working memory and the generation loop an unbounded-depth recurrence, which is why extra CoT tokens are compute, not decoration. Up to a point, letting the model reason longer raises accuracy: harder problems need more serial steps. But push too far and accuracy can sag — very long chains drift, second-guessing a correct interim answer or slipping on step forty a short chain never reached. This overthinking regime means the ideal chain length is problem-dependent, not ‘as long as possible.’

The practical shape is an inverted-U per difficulty band: easy problems peak at short chains and only lose accuracy (and money) if forced to ramble; hard problems keep improving far longer before their own turnover. A good reasoning model learns to modulate length — brief on the trivial, expansive on the genuinely hard — rather than applying a fixed long budget to everything. That length-control behavior is not something prompting reliably produces; it is largely something training installs, which is the bridge to the next section.