The limits are asymmetric, and that is the story

Everyone knows T → 0 gives the argmax and T → ∞ gives uniform. What matters is how fast each limit arrives, and the two rates are nothing alike. Let Δ = z_max − z_2nd. Cold, the runner-up’s mass decays exponentially: 1 − p_max ≈ exp(−Δ/T). Hot, expanding the softmax to second order in 1/T shows entropy approaches its ceiling only quadratically: ln V − H(T) ≈ Var_unif(z) / (2T^2), where Var_unif is the plain unweighted variance of the logits.

Check both on z = [5.0, 3.5, 3.0, 1.0, 0.5], where Δ = 1.5 and Var_unif = 2.74. At T = 0.3 the cold formula predicts 1 − p_max = 0.0067 against an actual 0.0079; at T = 3 the hot one predicts H = 1.457 nats against an actual 1.469. So cooling is a cliff and heating is a long ramp: T = 0.5 is already nearly greedy, while T = 3 is still far short of uniform. The usable band sits roughly in [0.5, 2] not by convention but because that is where the derivative actually lives.

Advertisement

The exchange rate: one nat costs T logits

The companion piece establishes dH/dT = Var_{p_T}(z) / T^3 ≥ 0. Diversity has a price, and the same exponential-family identities name it exactly. With β = 1/T we have d E_β[z] / dβ = Var_β(z), so the expected logit of the token you actually draw falls as you heat:

E(T)   = E_{p_T}[z]              expected logit of the sampled token
dE/dT  = Var_{p_T}(z) * (-1/T^2) = -Var / T^2
dH/dE  = (Var/T^3) / (-Var/T^2)  = -1/T      =>   dE = -T * dH

The variance cancels. Whatever the distribution’s shape, one nat of entropy costs exactly T logits of expected score — and because log p_1(i) = z_i − logZ(1) with logZ(1) constant within the step, that same −T is the price in expected model log-likelihood. It is the thermodynamic identity in a decoder’s clothes, and the honest form of the diversity-versus-coherence tradeoff: a slope you read off the T you are already running. At T = 0.5 a nat is cheap; at T = 2 the identical nat costs four times as much.

Advertisement

Two perplexities, and the gap between them

‘Perplexity’ at generation time names two different quantities. Distribution perplexity exp(H(p_T)) is the branching factor the sampler faces at that step. Text perplexity exp(−E_{p_T}[log p_1]) scores tokens drawn from p_T under the model’s own untempered distribution — the quality proxy an evaluator reports. Cross-entropy splits as entropy plus divergence:

−E_{p_T}[log p_1] = H(p_T) + KL(p_T || p_1)

So text perplexity is never below distribution perplexity, and they coincide only at T = 1. Same logits as above:

Texp H_1exp H_2text pplKL(p_T||p_1)
0.51.321.141.540.156
0.71.721.371.810.049
1.02.351.802.350.000
1.32.892.252.980.029
1.63.332.663.650.091
2.03.763.124.530.186

The KL column is a valley floored at T = 1: cooling departs from the model’s beliefs as surely as heating does. Greedy is not the safe setting, only the most confident-looking one.

Repetition is a Rényi-2 quantity, not a Shannon one

Entropy answers ‘how uncertain.’ Repetition asks something sharper: how often do two visits to the same state pick the same token? That is collision probability, and it is Rényi-2:

C        = Sigma_i p_i^2     probability two iid draws agree
H_2      = -log C            Renyi-2 entropy
exp(H_2) = 1 / C             effective number of choices

Because H_2 ≤ H_1 always, exp(H_2) ≤ exp(H_1): the collision count is the pessimistic number, and it is the one that tracks observed repeats. In the table, T = 1 offers 2.35 effective choices by Shannon but only 1.80 by collision — a 55% chance two passes agree. Two cautions. This is a per-step quantity, while the distinct-n metric from the NLG literature is a ratio of unique to total n-grams over text whose every step is conditioned on a different prefix — measure that one, do not derive it. And C is dominated by the head — a fattened tail barely moves it, which is why heating delivers less repetition relief than the entropy number promises.

The right T belongs to the checkpoint, not just the task

Task-dependence is covered thoroughly elsewhere in this series: extraction wants cold, brainstorming wants hot. The under-discussed half is that the same T means different things on different checkpoints, because T only rescales a sharpness the model already has. Base pretrained models are comparatively flat; instruction tuning and RLHF concentrate mass, distillation onto a teacher’s near-argmax targets concentrates it hard, and bigger, longer-trained models are more confident wherever they are confident at all. A 1B CPU-served SLM distilled from a larger teacher can sit most of a nat below a same-size base model at T = 1.

The consequence is blunt: a temperature is not portable. Copying T = 0.8 off a model card onto a different checkpoint transfers the wrong parameter, because the quantity you care about — the branching actually offered — lands somewhere else entirely. What does port is a target: ‘about 1.2 nats,’ or ‘effective support near 3.’ Then solve for the T that hits it on your own model.

Entropy-targeted and per-token temperature

That target is cheap to solve for, because the derivative is already in hand. Newton’s method on H(T) = H* steps by ΔT = (H* − H(T)) · T^3 / Var_{p_T}(z). On the running example with H* = 1.2 nats, starting from T = 1, the iterates are 1.000 → 1.409 → 1.572 → 1.592 — four decimals in four steps. H(T) is concave here, so the first step undershoots and the iterates climb monotonically; capping at three is safe, not just cheap.

Run that every token and you have an adaptive sampler: confident steps get a low T, uncertain steps a high one, and the branching factor is held constant instead of the multiplier. The cost is a handful of O(V) reductions — a few hundred thousand FLOPs against a decode step already spending billions, well under 0.1% even on a CPU SLM where every millisecond shows. Cruder schedules are worth trying first: decay T across the generation to open with variety and close with coherence, or key T to the top-1 minus top-2 logit margin.