The rule: a rank cut, then a renormalization

Given logits z: [V] and their softmax p, top-k keeps the survivor set S_k of the k highest entries, zeroes everything else, and rescales so the result is a distribution again:

S_k  = { i : rank(z_i) ≤ k }        rank descending
M_k  = Σ_{i ∈ S_k} p_i           retained mass
p’_i = p_i / M_k  if i ∈ S_k,   else 0

Two things deserve more attention than they usually get. First, S_k can be found on the raw logits: softmax is strictly increasing, so ranking by p and ranking by z give the same set. You never need to exponentiate the whole vocabulary. Second, M_k alone describes the damage — the divergence from truncating collapses to log(1/M_k), as the companion sampling article derives. Everything interesting about top-k is a statement about how M_k behaves when k is fixed and the context varies.

Advertisement

Rank is the one thing temperature cannot move

Dividing logits by T > 0 is strictly increasing, so it cannot reorder anything. Top-k therefore selects the identical k tokens at every temperature — but the mass those tokens carry is not identical at all. Take a 32k-vocabulary model with top logits [18.2, 14.6, 13.9, 13.1, 12.4, …] over a bulk near zero, held at k = 40:

Tp_(1)M_40log(1/M_40)
0.70.9910.99999~0.00001 nats
1.00.9460.99760.0024 nats
1.50.6040.75870.276 nats
2.00.1350.22071.511 nats

Same 40 tokens, same k, and the distortion moves five orders of magnitude. At T = 0.7 the setting is a no-op; at T = 2.0 it deletes 78% of the model’s own belief. The parameter did not change, and the thing it controls changed completely.

Advertisement

The order-statistics view of z_(k)

Write z_(1) ≥ z_(2) ≥ … ≥ z_(V) for the sorted logits. Then z_(k) is an order statistic, and top-k is really a threshold at z_(k) whose numeric value you never specify. If the vocabulary bulk is roughly N(μ, σ) — a decent first approximation away from the head — then the standard quantile estimate gives:

z_(k) ≈ μ + σ · Φ^-1(1 − k/V)

It is accurate: the step above has a N(0, 2) bulk and z_(40) = 6.03, against a predicted 0 + 3.02 · 2 = 6.04. For V = 32000: k = 10 → 3.42σ, 40 → 3.02σ, 100 → 2.73σ, 400 → 2.24σ. A 40× increase in k moves the cut by only 1.2σ, because for a Gaussian tail the threshold grows roughly like sqrt(2 ln(V/k)) — logarithmically. The same collapse shows in the spacings: in the step above, z_(1) − z_(2) = 3.6 logits but z_(40) − z_(100) = 0.50. Moving k from 1 to 2 is a violent change; moving it from 40 to 100 barely moves the boundary at all.

Failure one: the confident step, where k does nothing

Keep that step at T = 1. Its entropy is H = 0.31 nats, so the effective support exp(H) = 1.4 tokens; the leader holds 0.946 and four tokens already cover 99% of the mass. Now apply k = 40. The retained mass is M_40 = 0.9976, and the thirty tokens ranked 11 through 40 hold, after renormalization, a combined 2.7 × 10^-4 — roughly one draw in 3,700.

This is not a harmful setting, it is an inert one. Top-k protected you from nothing, because the model had already crushed its own tail; you paid a selection pass over 32,000 logits to move the sampling distribution by a quarter of a percent. The failure is not bad output but false confidence: clean text on confident steps gets credited to k = 40 when the softmax did all of it.

Failure two: the open step, where k cuts the model off

Now a genuinely open step — the same model mid-narrative, where about a hundred continuations are live. Entropy H = 4.46 nats, effective support exp(H) = 86, leader at p_(1) = 0.074, and it takes 80 tokens to accumulate 0.90 of the mass.

Apply the identical k = 40: M_40 = 0.690. You have deleted 31% of the model’s belief, log(1/0.690) = 0.371 nats of distortion — 155× the distortion the same setting imposed on the confident step. The first dropped token, which the model rated at p_(41) = 0.0072, is now strictly impossible, and over a long generation those unreachable-but-plausible continuations compound into a measurably narrower voice — at precisely the steps where diversity was the point. Fixed k is loosest where the model is certain and tightest where it is not: exactly backwards.

Retained mass is a random variable

Put the two steps side by side and the structural problem is plain: with k pinned at 40, M_k took the values 0.9976 and 0.690 within a single generation from a single model. M_k is not a constant you configured — it is a random variable over contexts, and k only fixes its index, never its value.

That reframes tuning k. You are not choosing a truncation strength; you are choosing a rank and letting the next token’s entropy choose the strength for you. Per-token entropy in real text swings by an order of magnitude — near-deterministic inside a word or a known proper noun, wide open after a sentence boundary — so the induced spread of M_k is enormous. Mass-based and peak-relative rules exist to shrink that spread by construction; each has its own article in this series.

Selection: you do not need a sort

The naive implementation sorts the whole vocabulary, O(V log V), then slices. That is more work than the problem requires, because renormalization does not care about order — you need the survivor set, not a ranking of it.

full sort        O(V log V)   sorts 32,000 to use 40
quickselect      O(V) expected, in place, unordered output
bounded heap     O(V log k)   one pass, k-sized working set

A single-threaded NumPy measurement on a 32,000-element float32 vector: full sort 156 µs, argpartition for k = 40 73 µs — a bit over 2×. The larger win is upstream: because S_k is decidable on raw logits, a top-k-only pipeline never exponentiates the vocabulary. Select in O(V) comparisons, then exp over k values — 40 calls, not 32,000.