The rule: a rank cut, then a renormalization
Given logits z: [V] and their softmax p, top-k keeps the survivor set S_k of the k highest entries, zeroes everything else, and rescales so the result is a distribution again:
S_k = { i : rank(z_i) ≤ k } rank descending
M_k = Σ_{i ∈ S_k} p_i retained mass
p’_i = p_i / M_k if i ∈ S_k, else 0Two things deserve more attention than they usually get. First, S_k can be found on the raw logits: softmax is strictly increasing, so ranking by p and ranking by z give the same set. You never need to exponentiate the whole vocabulary. Second, M_k alone describes the damage — the divergence from truncating collapses to log(1/M_k), as the companion sampling article derives. Everything interesting about top-k is a statement about how M_k behaves when k is fixed and the context varies.
Rank is the one thing temperature cannot move
Dividing logits by T > 0 is strictly increasing, so it cannot reorder anything. Top-k therefore selects the identical k tokens at every temperature — but the mass those tokens carry is not identical at all. Take a 32k-vocabulary model with top logits [18.2, 14.6, 13.9, 13.1, 12.4, …] over a bulk near zero, held at k = 40:
| T | p_(1) | M_40 | log(1/M_40) |
|---|---|---|---|
| 0.7 | 0.991 | 0.99999 | ~0.00001 nats |
| 1.0 | 0.946 | 0.9976 | 0.0024 nats |
| 1.5 | 0.604 | 0.7587 | 0.276 nats |
| 2.0 | 0.135 | 0.2207 | 1.511 nats |
Same 40 tokens, same k, and the distortion moves five orders of magnitude. At T = 0.7 the setting is a no-op; at T = 2.0 it deletes 78% of the model’s own belief. The parameter did not change, and the thing it controls changed completely.
The order-statistics view of z_(k)
Write z_(1) ≥ z_(2) ≥ … ≥ z_(V) for the sorted logits. Then z_(k) is an order statistic, and top-k is really a threshold at z_(k) whose numeric value you never specify. If the vocabulary bulk is roughly N(μ, σ) — a decent first approximation away from the head — then the standard quantile estimate gives:
z_(k) ≈ μ + σ · Φ^-1(1 − k/V)It is accurate: the step above has a N(0, 2) bulk and z_(40) = 6.03, against a predicted 0 + 3.02 · 2 = 6.04. For V = 32000: k = 10 → 3.42σ, 40 → 3.02σ, 100 → 2.73σ, 400 → 2.24σ. A 40× increase in k moves the cut by only 1.2σ, because for a Gaussian tail the threshold grows roughly like sqrt(2 ln(V/k)) — logarithmically. The same collapse shows in the spacings: in the step above, z_(1) − z_(2) = 3.6 logits but z_(40) − z_(100) = 0.50. Moving k from 1 to 2 is a violent change; moving it from 40 to 100 barely moves the boundary at all.
Failure one: the confident step, where k does nothing
Keep that step at T = 1. Its entropy is H = 0.31 nats, so the effective support exp(H) = 1.4 tokens; the leader holds 0.946 and four tokens already cover 99% of the mass. Now apply k = 40. The retained mass is M_40 = 0.9976, and the thirty tokens ranked 11 through 40 hold, after renormalization, a combined 2.7 × 10^-4 — roughly one draw in 3,700.
This is not a harmful setting, it is an inert one. Top-k protected you from nothing, because the model had already crushed its own tail; you paid a selection pass over 32,000 logits to move the sampling distribution by a quarter of a percent. The failure is not bad output but false confidence: clean text on confident steps gets credited to k = 40 when the softmax did all of it.
Failure two: the open step, where k cuts the model off
Now a genuinely open step — the same model mid-narrative, where about a hundred continuations are live. Entropy H = 4.46 nats, effective support exp(H) = 86, leader at p_(1) = 0.074, and it takes 80 tokens to accumulate 0.90 of the mass.
Apply the identical k = 40: M_40 = 0.690. You have deleted 31% of the model’s belief, log(1/0.690) = 0.371 nats of distortion — 155× the distortion the same setting imposed on the confident step. The first dropped token, which the model rated at p_(41) = 0.0072, is now strictly impossible, and over a long generation those unreachable-but-plausible continuations compound into a measurably narrower voice — at precisely the steps where diversity was the point. Fixed k is loosest where the model is certain and tightest where it is not: exactly backwards.
Retained mass is a random variable
Put the two steps side by side and the structural problem is plain: with k pinned at 40, M_k took the values 0.9976 and 0.690 within a single generation from a single model. M_k is not a constant you configured — it is a random variable over contexts, and k only fixes its index, never its value.
That reframes tuning k. You are not choosing a truncation strength; you are choosing a rank and letting the next token’s entropy choose the strength for you. Per-token entropy in real text swings by an order of magnitude — near-deterministic inside a word or a known proper noun, wide open after a sentence boundary — so the induced spread of M_k is enormous. Mass-based and peak-relative rules exist to shrink that spread by construction; each has its own article in this series.
Selection: you do not need a sort
The naive implementation sorts the whole vocabulary, O(V log V), then slices. That is more work than the problem requires, because renormalization does not care about order — you need the survivor set, not a ranking of it.
full sort O(V log V) sorts 32,000 to use 40
quickselect O(V) expected, in place, unordered output
bounded heap O(V log k) one pass, k-sized working setA single-threaded NumPy measurement on a 32,000-element float32 vector: full sort 156 µs, argpartition for k = 40 73 µs — a bit over 2×. The larger win is upstream: because S_k is decidable on raw logits, a top-k-only pipeline never exponentiates the vocabulary. Select in O(V) comparisons, then exp over k values — 40 calls, not 32,000.