Self-consistency is the simplest test-time compute technique that reliably works: ask a model the same reasoning question several times with sampling turned on, pull the final answer out of each response and return the most common one. Correct reasoning paths tend to converge on the same answer while mistakes scatter, so the plurality is right more often than any single path. Wang and colleagues introduced it in 2022 and reported large gains over single chain-of-thought on arithmetic and commonsense benchmarks.

The idea fits in a sentence, and the self-consistency overview explains why voting works. This article is about making it work in a real system: extracting and normalising answers so votes are counted correctly, choosing the number of samples, stopping early, extending it to free-form outputs, using the vote margin as a confidence score and proving the gain on your own data before paying for it.

Advertisement

The pipeline

A self-consistency call has five stages: a prompt that elicits reasoning and a fixed answer format, N sampled completions, extraction of the final answer from each, normalisation so equivalent answers compare equal, and a vote. In production add a sixth: a decision on the result based on how strongly the samples agreed.

Self-consistency as a pipeline: the vote is the easy partPromptCoT + answer formatSample 1temperature 0.7Sample 2Sample 3Sample N...Extract andnormalise answersVotepluralityAccepthigh agreementEscalatelow agreementEarly stopleader cannot be overtakenCost grows with N; extraction errors silently split the vote
Self-consistency as a pipeline. Samples run in parallel; extraction and normalisation turn free text into comparable votes; agreement decides whether to accept or escalate; early stopping saves samples once the winner is settled.

Self-consistency only applies where answers can be compared: a number, a choice, a label, a short entity, a SQL result set or the output of running code. If the output is an essay, plain voting has nothing to count, and you need the universal variant described below.

Extraction and normalisation decide everything

Voting fails silently when extraction is sloppy. If one sample ends with “The answer is 12”, another with “12 apples” and a third with “$12.00”, a naive string vote sees three different answers and the true consensus is lost. Worse, a regex that grabs the first number in a response will vote for an intermediate result.

Fix it at both ends. In the prompt, require a terminal line in a fixed format, such as Final answer: <value>, and give one example of it. In code, take the last matching line, normalise by answer type and treat an unparseable response as an abstention, never as a vote:

import re
from fractions import Fraction

ANSWER = re.compile(r"(?im)^\s*final answer\s*:\s*(.+?)\s*$")

def extract(completion):
    """Return the last 'Final answer:' line, or None (an abstention, not a vote)."""
    hits = ANSWER.findall(completion)
    return hits[-1] if hits else None

def normalise(ans, kind):
    if ans is None:
        return None
    s = ans.strip().rstrip(".").replace(",", "").replace("$", "")
    if kind == "number":
        lead = re.match(r"-?[\d.]+(?:/\d+)?", s)       # "12 apples" -> "12"
        s = lead.group(0) if lead else s
        try:
            return str(Fraction(s).limit_denominator(10_000))   # "0.5", "1/2" -> "1/2"
        except (ValueError, ZeroDivisionError):
            return None
    if kind == "choice":
        m = re.match(r"\(?([A-E])\)?\b", s.upper())
        return m.group(1) if m else None
    return " ".join(s.lower().split())

Normalising numbers through Fraction makes 0.5, 1/2 and .50 one vote. For multiple choice, map to a letter. For code, compare by behaviour: run each candidate against a few generated inputs and vote on the output signature, because two different programs can both be right. For SQL, vote on the result set. Log the raw answers that failed to parse; a rising abstention rate is the first sign a prompt or model change broke your format.

Advertisement

A parallel implementation with early stopping

Samples are independent, so issue them concurrently. The function below takes any async generate(prompt, temperature) that wraps your provider, samples in batches and stops as soon as the leader's margin exceeds the number of samples left, at which point no outcome of the remaining calls could change the winner:

import asyncio
from collections import Counter

async def self_consistent(generate, prompt, kind, n=10, temperature=0.7, batch=5,
                          stop_margin=True):
    """generate(prompt, temperature) -> completion text; any provider."""
    votes, abstain, used = Counter(), 0, 0
    while used < n:
        k = min(batch, n - used)
        outs = await asyncio.gather(*(generate(prompt, temperature) for _ in range(k)))
        used += k
        for out in outs:
            a = normalise(extract(out), kind)
            if a is None:
                abstain += 1
            else:
                votes[a] += 1
        if stop_margin and votes:
            ranked = votes.most_common(2)
            lead = ranked[0][1] - (ranked[1][1] if len(ranked) > 1 else 0)
            if lead > n - used:          # remaining samples cannot change the winner
                break
    if not votes:
        return None, 0.0, used
    answer, count = votes.most_common(1)[0]
    return answer, count / used, used    # agreement counts abstentions against you

Trace it with N of 10 and batches of 5. The first batch returns 42, 42, 42, 40 and one response with no final-answer line. Votes are 42 three times and 40 once, with one abstention; the margin is 2 and 5 samples remain, so sampling continues. The second batch returns 42, 42, 42, 38 and 42. Now 42 has 7 votes, the margin is 6 and no samples remain, so the loop ends with 42 at an agreement of 7 out of 10. Had the first batch been five clean 42s, the margin of 5 would only equal the 5 remaining samples, which still allows a tie, so the rule runs the second batch; with an odd budget of 9, the same five agreeing samples exceed the 4 remaining and the call stops at just over half its budget, which is what happens on most easy questions.

Agreement is computed over all samples used, including abstentions, so a run where half the responses were unparseable cannot report high confidence. If your provider can return several completions for one request, use that instead of separate calls, since the prompt is processed once. Either way, cache the shared prompt prefix where your provider supports prompt caching, because every sample repeats it.

Choosing N and temperature

Gains are steep for the first few samples and flatten quickly. A useful mental model: if each sample is right with probability 0.6 and wrong answers are spread across several values, a plurality of 5 is right noticeably more often than one sample, a plurality of 15 more still, and beyond that each extra sample buys little. If wrong answers concentrate on one value, typically a tempting trap, voting helps much less and can even lock in the error, because the trap wins the vote. That is why measured curves differ by task and you must measure your own.

Temperature controls diversity. Too low and the samples are near copies, so the vote adds nothing; too high and reasoning quality drops for every sample. Values around 0.5 to 0.8, with the provider's default top-p, are a reasonable starting range; sweep them on your evaluation set. Some models and reasoning modes fix or restrict sampling parameters, so check your provider's documentation rather than assuming a temperature setting is honoured.

Cost scales linearly in output tokens. With a 600-token reasoning chain and N of 10, one question costs 6,000 output tokens instead of 600. Early stopping cuts this sharply on easy questions, where the first five samples usually agree, and spends the full budget only on hard ones. Adaptive-Consistency (Aggarwal and colleagues, 2023) and Early-Stopping Self-Consistency (Li and colleagues, 2024) formalise this with statistical stopping rules and report large sample savings at similar accuracy; the margin rule above is the simplest exact version.

Weighted and universal variants

Plain plurality treats every sample equally. Three refinements are worth knowing.

VariantHow it worksUse when
Weighted by likelihoodEach vote weighted by the sample's normalised log-probabilityYour API returns token log-probs; gains are usually small
Verifier-weightedA separate scorer rates each reasoning path; votes weighted by scoreYou have or can train a reliable verifier
Universal self-consistencyThe model is shown all candidates and asked which is most consistent with the othersFree-form outputs such as summaries or open answers
Semantic clusteringEmbed answers, cluster near-duplicates, vote on clustersShort free-text answers with paraphrase variation

Universal self-consistency (Chen and colleagues, 2023) is the practical route for outputs that cannot be normalised: instead of counting, a final call selects the candidate that agrees most with the rest. It costs one more call and inherits that call's biases, such as preferring longer responses, so evaluate it rather than assuming it beats a single good sample. Verifier weighting is covered in the verifier article.

Agreement as a confidence signal

The most underused output of self-consistency is the vote share. When 9 of 10 samples agree, the answer is usually right; when the top answer holds 3 of 10, it often is not. That makes agreement a cheap confidence estimate for routing: accept high-agreement answers automatically, send low-agreement ones to a stronger model, retrieval, a tool call or a human.

Worked example. A support team extracts refund amounts from customer emails. On a labelled set of 500 emails, a single sample is right 88 percent of the time. Self-consistency at N of 7 lifts that to 93 percent. More useful, answers with agreement of at least 6 of 7 cover 81 percent of emails and are right 99 percent of the time; the other 19 percent go to a human queue. The automation now has a measured error rate instead of a hopeful one. These numbers are illustrative; the point is that the agreement threshold turns accuracy into a coverage-versus-precision dial you can set.

Calibrate the threshold on held-out data and recheck it whenever the model or prompt changes, because agreement levels shift with both.

Measuring it before paying for it

Never ship self-consistency on faith. Run the same labelled set at several values of N and record accuracy, calls per item, and coverage and precision at your acceptance threshold:

async def evaluate(dataset, generate, kind, ns=(1, 5, 10, 20)):
    for n in ns:
        correct = calls = accepted = acc_correct = 0
        for item in dataset:
            ans, agree, used = await self_consistent(generate, item["prompt"], kind, n=n)
            calls += used
            ok = ans == normalise(item["gold"], kind)
            correct += ok
            if agree >= 0.6:
                accepted += 1
                acc_correct += ok
        print(f"N={n:>2} acc={correct/len(dataset):.3f} calls/item={calls/len(dataset):.1f} "
              f"coverage@0.6={accepted/len(dataset):.2f} "
              f"acc_when_accepted={acc_correct/max(accepted,1):.3f}")

Compare against cheaper alternatives at equal cost: a stronger model with one sample, a better prompt, or retrieval. For models that already reason at length internally, the marginal gain from external voting is often smaller and the per-sample cost larger, so the comparison matters more. Integrate the harness with your existing prompt evaluation setup so regressions are caught in CI.

Failure modes

SymptomCauseFix
No gain over one sampleTemperature too low, samples identicalRaise temperature; check answer diversity
Consensus on a wrong answerSystematic error or trap shared by all pathsBetter prompt, retrieval or tools; voting cannot fix bias
Votes split across equivalent answersWeak normalisationType-aware normalisation; log raw answers
Confident answer from mostly broken outputAbstentions ignored in agreementCount abstentions in the denominator
Cost blows upFixed N on easy questionsEarly stopping; batch sampling; prompt caching
Latency spikesSequential sampling or one slow callParallel calls with a timeout; vote on what returned

Where it fits

Self-consistency sits between single chain-of-thought and search methods such as tree of thoughts. It is embarrassingly parallel, needs no change to the model and composes with almost anything: retrieval, tool use and verifiers. Use it when answers are checkable, errors are varied rather than systematic and the cost of a wrong answer exceeds a few extra model calls.

What to do next

  1. Pick one task with a checkable answer and build a labelled set of at least 200 items.
  2. Add a fixed final-answer line to the prompt and a type-aware extractor that abstains on parse failure.
  3. Run the evaluation harness at N of 1, 5, 10 and 20 and plot accuracy against calls per item.
  4. If the curve justifies it, deploy with parallel batched sampling and margin-based early stopping.
  5. Choose an agreement threshold from held-out data and route low-agreement answers to escalation.
  6. Monitor abstention rate, average samples used and agreement distribution, and re-run the harness after every model or prompt change.
Key takeaway: Self-consistency samples several reasoning paths and returns the most common final answer, and it reliably improves accuracy when answers are checkable and errors are varied. In production the vote is the easy part: enforce a final-answer format, normalise answers by type, count unparseable outputs as abstentions, sample in parallel and stop once the winner is settled. Use the agreement share as a confidence score to accept or escalate, and measure accuracy against cost on your own labelled data before deploying it.