Scoring as a Bernoulli process
Start with the simplest honest model of a benchmark. Each item is graded correct or incorrect, so item i is a Bernoulli trial x_i ∈ {0, 1} that comes up 1 with some true probability p — the model’s real competence on that population of tasks. The reported accuracy is just the sample mean:
p_hat = (1/n) Σ_i x_i x_i ∈ {0, 1}
E[p_hat] = p (unbiased)
Var(p_hat) = p(1 - p) / nTwo facts fall out immediately. The estimate is unbiased: on average it lands on the true p. And its variance shrinks like 1/n, so precision improves only with the square root of the number of items — quadrupling the test set halves the error bar. That p(1-p) factor also says variance is largest near p = 0.5 and smallest near 0 or 1, which is why mid-range scores are the noisiest to pin down.
The standard error hiding behind every score
The standard deviation of the accuracy estimate is its standard error, and because we don’t know the true p we plug in p_hat:
SE = √( p_hat (1 - p_hat) / n )Work a concrete case. A model scores 72% on n = 200 items. Then SE = √(0.72 × 0.28 / 200) ≈ √(0.001008) ≈ 0.0317, about 3.2 percentage points. A 95% interval spans roughly ±1.96 × SE ≈ ±6.2 points — so ‘72%’ really means ‘somewhere around 66–78%.’ Any two models within that band on this test are statistically indistinguishable, no matter how confidently a table ranks one above the other. Reporting a bare accuracy without its n and error bar is the single most common way eval numbers mislead.
Confidence intervals: Wald, Wilson, Clopper-Pearson
The p_hat ± z·SE interval above is the Wald interval, and it is comfortable but flawed: near p = 0 or 1, or for small n, it produces impossible bounds below 0 or above 1 and under-covers badly. Prefer the Wilson score interval, which stays inside [0, 1] and behaves well for extreme rates and small samples:
center = (p_hat + z²/2n) / (1 + z²/n)
half = ( z √( p_hat(1-p_hat)/n + z²/4n² ) ) / (1 + z²/n)When you need a guaranteed-coverage exact interval — a model that solved 10/10 items, say, where Wald degenerates to zero width — use the Clopper-Pearson interval, computed from binomial tail probabilities. It is conservative (slightly too wide) but never lies. As a rule: Wilson for everyday reporting, Clopper-Pearson when a hard guarantee matters.
How many items do you actually need?
Sample size is a planning question you can answer before spending a dollar of inference. To hold the half-width of a 95% interval to a margin m, invert the standard error:
n ≥ z² · p(1-p) / m²
worst case p = 0.5: n ≥ z² / (4 m²)For a ±3 point margin at 95% confidence (z = 1.96, m = 0.03), the worst-case requirement is n ≥ 1.96² / (4 × 0.03²) ≈ 1068 items. Want ±1 point? That balloons to about 9600 items, because the 1/m² term punishes precision quadratically. This is why most public benchmarks — a few hundred questions — simply cannot resolve small differences, and why a serious evaluation either buys more items or accepts that sub-margin gaps are noise.
pass@k and the math of multiple samples
Many capability tests let the model try k times and count success if any attempt passes — the pass@k metric. Estimating it naively (generate k, check) is high variance, so the standard trick from the HumanEval work is to generate a larger pool of n ≥ k samples, count c correct, and use the unbiased combinatorial estimator:
pass@k = 1 − C(n − c, k) / C(n, k)The fraction is the probability that a random draw of k from the pool lands entirely among the wrong answers; one minus that is the chance at least one is right. Compute it in log-space to avoid overflow in the binomials. The lesson is that pass@k rises with k for free — more attempts buy more coverage — so comparing a pass@1 number against someone else’s pass@10 is comparing different quantities, not different models.
Comparing two models without fooling yourself
Ranking two models is not ‘whose accuracy is higher’ but ‘is the gap larger than the noise.’ When both models answer the same items, the comparison is paired, and the right tool is McNemar’s test, which looks only at the items where they disagree:
b = A right, B wrong c = A wrong, B right
χ² = (|b − c| − 1)² / (b + c)Items both got right or both got wrong carry no information about the difference and are correctly ignored. If b + c is small, even a lopsided split isn’t significant — you simply don’t have enough disagreements to tell. For unpaired setups use a two-proportion z-test, but pairing is strictly more powerful and almost always available in benchmarking, so prefer it. Never declare a winner from raw accuracies whose confidence intervals overlap.
Item response theory: not every question is equal
Accuracy treats all items as interchangeable, but a benchmark mixes trivial and brutal questions. Item response theory (IRT) models the probability a model of latent ability θ answers an item correctly as a logistic curve in the item’s parameters — the two-parameter (2PL) model:
P(correct | θ) = 1 / (1 + exp( −a (θ − b) ))Here b is item difficulty (the ability at which success is 50/50) and a is discrimination (how sharply the item separates strong from weak models). Setting a = 1 recovers the one-parameter Rasch model. Instead of one raw percentage, you fit each model an ability θ on a common scale and each item its own difficulty, so scores from different item mixes become comparable — and easy questions stop inflating a weak model’s number. Difficulties also enable adaptive testing (serve items near a model’s θ to reach precision in fewer questions), expose benchmark saturation when every item sits below the frontier, and let low discrimination flag junk or contaminated items.
LLM-as-judge and measurement noise
When answers are open-ended, a model often grades them, and the judge is itself a noisy, biased instrument. Formally the observed score is the true label plus a judge error term: even a good judge flips some verdicts, and a biased one systematically favours longer answers, its own family’s style, or the first option shown. Quantify reliability before trusting the numbers: measure judge–human agreement with Cohen’s kappa, which corrects raw agreement for chance:
κ = (p_observed − p_chance) / (1 − p_chance)A κ around 0.8 is solid; near 0.4 the judge is barely better than a coin and its scores can’t separate close models. Mitigations include position-swapping to cancel order bias and averaging several judges. Judge noise adds directly to sampling noise, so an under-validated judge can manufacture or erase a leaderboard gap on its own.
Contamination: the silent score inflator
Every statistic above assumes the model has not seen the test. When benchmark items leak into pretraining data — and with web-scraped corpora they routinely do — the model can recall answers instead of deriving them, and the score measures memorisation, not capability. Contamination is the most dangerous eval failure precisely because it looks like success: the numbers go up, the error bars stay tight, and nothing in the accuracy itself reveals the leak. It inflates public benchmarks hardest, since those are the pages most likely to be crawled and quoted across the internet. Any capability claim on a well-known public test set should be read with contamination as the default suspicion — the burden is on the evaluator to show the model hadn’t already read the answer key.