From logits to a distribution

Each decode step produces z: [V], one real score per vocabulary entry, with V typically 32k–256k. The softmax turns it into a distribution:

p_i = exp(z_i) / Σ_j exp(z_j)

One property governs everything downstream: softmax is shift invariant. Adding the same constant to every logit changes nothing, because exp(z_i + c) / Σ_j exp(z_j + c) = exp(z_i) / Σ_j exp(z_j). Logits therefore carry no absolute meaning — only differences are real. (Implementations exploit this by subtracting max(z) before exponentiating, purely for numerical stability.)

Keep that in hand, because it sorts the sampling knobs into three honest classes. Temperature rescales the differences. Truncation deletes entries. Penalties shift individual entries. Any knob whose effect depends on where zero happens to sit is, as we will see, ill-posed.

Advertisement

Temperature: divide before you exponentiate

Temperature T > 0 divides the logits before the softmax:

p_i(T) = exp(z_i / T) / Σ_j exp(z_j / T)
p_i(T) / p_j(T) = exp( (z_i − z_j) / T )

The second line is the whole mechanism. Temperature rescales every log-odds ratio by 1/T. Halving T squares the odds between any pair of tokens; doubling it takes their square root. The ranking never changes — division by a positive scalar is monotone — only the contrast does. Limits: T → 0+ concentrates all mass on the argmax, T → ∞ approaches uniform over V.

Take logits z = [4, 3, 2, 0]. At T = 1 the probabilities are [.657, .242, .089, .012]; at T = 0.5, [.867, .117, .016, .0003]; at T = 2, [.474, .288, .174, .064]. Notice the last entry: it moves from 0.03% to 6.4%, a factor of over 200. Temperature does its most dramatic work in the tail — and a real vocabulary has tens of thousands of tail entries, not one.

Advertisement

Why temperature is exactly an entropy dial

“Higher temperature means more randomness” is usually asserted. It is provable. Write β = 1/T and let logZ(β) = log Σ_j exp(β z_j). Standard exponential-family identities give d logZ / dβ = E[z] and d E[z] / dβ = Var(z), both under p_β. Since H = logZ − β E[z]:

dH/dβ = E[z] − E[z] − β · Var(z) = −β · Var(z)
dH/dT     = (dH/dβ)(dβ/dT) = (−β Var(z))(−1/T²) = Var_{p_T}(z) / T³  ≥ 0

Entropy is monotonically non-decreasing in T, strictly increasing unless every logit is equal, and the rate is the variance of the logits under the current distribution. Note Var_{p_T}(z) is itself a function of T, so this is not a 1/T³ growth law — the variance shrinks as the distribution flattens.

For z = [4, 3, 2, 0]: H = 0.44 / 0.89 / 1.19 nats at T = 0.5 / 1 / 2, against a ceiling of ln 4 = 1.39. Temperature is the only knob here that is a smooth, principled control on uncertainty.

The truncation family: three answers to one question

Truncation asks: which tokens keep nonzero probability? Every method picks a survivor set S, zeroes the rest, and renormalizes over S. The three standard answers differ only in how S is chosen.

Top-k fixes the count: keep the k highest-probability tokens, always exactly k of them, regardless of whether the model is certain or lost. Top-p (nucleus) fixes the mass: sort descending and keep the shortest prefix whose cumulative probability reaches p. Its count adapts — one token when the model is confident, hundreds when it is not. Both have dedicated articles in this series with their full derivations; here they matter as family members.

Min-p fixes the relative floor: keep every token with p_i ≥ m · max_j p_j. It needs neither a sort nor a cumulative sum, and it is the member whose interaction with temperature is most worth deriving — which is the next section.

Min-p is a window in logit space

Substitute the softmax into the min-p test and the sums cancel:

p_i ≥ m · p_max
  ⇔ exp((z_i − z_max)/T) ≥ m
  ⇔ z_i ≥ z_max − T · ln(1/m)

So min-p is not really a probability threshold at all: it keeps every token inside a window of width T · ln(1/m) below the top logit. At m = 0.05 that is 3.0 logits wide at T = 1 and 4.5 at T = 1.5. This is why min-p is described as confidence-adaptive: a peaked distribution has few tokens in the window, a flat one has many, with no sorting required.

Two consequences. First, if your stack applies min-p to the post-temperature distribution, the window widens linearly with T — raising temperature quietly loosens your truncation, so check where your backend puts it. Second, the cost: min-p is a max pass plus a threshold pass, O(V) with no sort, versus O(V log V) for a naive top-p. On a CPU SLM emitting 30 tokens/second against a 128k vocabulary, that difference is measurable step time, not a rounding error.

Penalties edit history — and only additive ones are well-posed

Penalties reach outside the current step and modify logits using what has already been generated. The two additive forms, over the seen-token counts c_i:

presence:  z_i ← z_i − α · [c_i > 0]
frequency: z_i ← z_i − γ · c_i

Subtracting α from a logit multiplies that token’s odds against every other by exp(−α): 0.61× at α = 0.5, 0.37× at 1.0, 0.14× at 2.0. Presence is a one-time toll; frequency compounds without bound, so a token seen twenty times is effectively banned — fine for a stuck phrase, disastrous for the.

The classic repetition penalty is multiplicative instead: z_i ← z_i / r when z_i > 0, z_i · r otherwise. This breaks the shift invariance from section one. softmax(z + c) = softmax(z), but softmax((z + c)/r) ≠ softmax(z/r) — so its strength depends on whether the stack hands it raw logits, max-subtracted logits, or log-probabilities. Same number, different behaviour. Prefer the additive penalties when you have the choice.