Every API that generates text exposes a few numbers that most people copy from an example and never touch again: temperature, top_p and sometimes top_k. They do not change what the model knows. They change how one token is picked from the probability distribution the model produces at each step, and over hundreds of steps that choice decides whether an answer is consistent, varied or incoherent.
This page explains the knobs from first principles with one worked example, shows a sampler you can read in thirty lines, lists what the main APIs actually accept, and ends with a method for choosing settings by measurement instead of folklore. The engine-side architecture, including penalties and the logits pipeline in a serving stack, is covered in LLM sampling and decoding architecture; this page is about choosing and tuning the settings as the person writing the prompt.
What the model hands the sampler
At each step the model outputs one score, a logit, for every token in its vocabulary. Logits are unnormalised: only their differences matter. The softmax turns them into probabilities: pi = exp(zi) / Σj exp(zj). The sampler then draws one token, appends it to the context and the model runs again.
Take a prompt ending in The capital of France is and suppose the six largest logits are Paris 5.0, the 3.2, a 2.9, London 2.0, Lyon 1.5 and banana -1.0. The rest of the vocabulary is ignored here to keep the arithmetic visible. At the default setting the softmax gives Paris 0.730, the 0.121, a 0.089, London 0.036, Lyon 0.022 and banana 0.002. Greedy decoding always picks Paris. Plain sampling picks Paris 73% of the time, and London about once in 28 draws, which over a 500-token answer means a wrong turn is almost certain somewhere. Every knob below exists to control that tail.
Temperature reshapes the whole distribution
Temperature T divides every logit before the softmax: pi = softmax(zi / T). Dividing by a number below 1 stretches the gaps between logits, so the leader gains probability; dividing by a number above 1 shrinks the gaps, so the distribution flattens. T = 1 leaves the model's distribution unchanged. As T approaches 0 the result approaches greedy argmax, and implementations special-case T = 0 as argmax rather than dividing by zero.
| T | Paris | the | a | London | Lyon | banana | Entropy (bits) |
|---|---|---|---|---|---|---|---|
| 0.5 | 0.956 | 0.026 | 0.014 | 0.002 | 0.001 | 0.000 | 0.32 |
| 0.7 | 0.872 | 0.067 | 0.043 | 0.012 | 0.006 | 0.000 | 0.75 |
| 1.0 | 0.730 | 0.121 | 0.089 | 0.036 | 0.022 | 0.002 | 1.32 |
| 1.3 | 0.615 | 0.154 | 0.122 | 0.061 | 0.042 | 0.006 | 1.70 |
| 2.0 | 0.454 | 0.185 | 0.159 | 0.101 | 0.079 | 0.023 | 2.14 |
The numbers were computed with a short script from the six logits above. Two things stand out. Temperature never removes a token: banana is still possible at T = 0.5, just rare. And the effect is multiplicative in log space, so it hits the tail hardest: London falls from 0.036 to 0.002 between T = 1 and T = 0.5, about a 15-fold drop before rounding, while Paris rises by less than a third. Entropy is a convenient single number for how spread out the choice is; it roughly doubles from T = 0.7 to T = 1.3 here.
Top-k, top-p and min-p cut the tail off
Truncation samplers set the probability of unlikely tokens to zero and renormalise the rest. They differ in how they decide what is unlikely.
- Top-k keeps the k highest-probability tokens. It is blind to shape: with k = 40, a confident step keeps 39 junk tokens, and a genuinely open step, say choosing the next word of a story, may cut real options.
- Top-p, or nucleus sampling, keeps the smallest set of tokens whose cumulative probability reaches p. At T = 1 with p = 0.9 our example keeps Paris, the and a (0.730 + 0.121 + 0.089 = 0.940) and drops London, Lyon and banana. The set grows on flat steps and shrinks on confident ones, which is why top-p replaced top-k as the usual default.
- Min-p keeps tokens whose probability is at least m times the top token's. With m = 0.05 the threshold here is 0.05 × 0.730 = 0.0365, so London at 0.036 just misses and the same three tokens survive. Min-p scales with the model's confidence directly: on a confident step the cutoff is high and almost everything is dropped, on an open step the cutoff falls and more options survive. Where it runs relative to temperature matters, as the next section shows; with temperature applied first, a hot setting lowers the top probability and lets more tokens through.
None of these needs to be combined with the others, and stacking all three mostly makes the configuration hard to reason about. Pick one truncation method and one temperature, then measure.
Order matters more than most documentation admits
If temperature is applied before truncation, a high temperature flattens the distribution first and more tokens then survive top-p. If truncation runs first, the surviving set is decided on the unscaled distribution and temperature only reweights the survivors. With our logits at T = 2, temperature-first top-p 0.9 keeps five tokens including Lyon. Truncate-first keeps the same three it keeps at T = 1 and reweights them to Paris 0.569, the 0.231 and a 0.199.
llama.cpp documents its default chain as penalties;dry;top_n_sigma;top_k;typ_p;top_p;min_p;xtc;temperature, with temperature last, and lets you reorder it with --samplers. Other engines and hosted APIs do not always document their order, and some apply temperature first. When you move a prompt between engines, the same numbers can produce visibly different behaviour, so re-tune rather than copy settings across.
A reference sampler you can read
This is a complete, temperature-first sampler in NumPy. It is not fast, but every production sampler is a faster version of these steps, and it is useful for checking what a configuration does to a logit vector you have captured.
import numpy as np
def sample(logits, temperature=1.0, top_k=0, top_p=1.0, min_p=0.0, rng=None):
rng = rng or np.random.default_rng()
z = np.asarray(logits, dtype=np.float64)
if temperature <= 0: # special case: greedy
return int(np.argmax(z))
z = z / temperature
p = np.exp(z - z.max()); p /= p.sum() # stable softmax
keep = np.ones_like(p, dtype=bool)
if top_k > 0:
kth = np.sort(p)[-min(top_k, p.size)]
keep &= p >= kth
if top_p < 1.0:
order = np.argsort(-p)
cum = np.cumsum(p[order])
cutoff = np.searchsorted(cum, top_p) + 1 # smallest prefix reaching top_p
mask = np.zeros_like(keep); mask[order[:cutoff]] = True
keep &= mask
if min_p > 0.0:
keep &= p >= min_p * p.max()
p = np.where(keep, p, 0.0); p /= p.sum()
return int(rng.choice(p.size, p=p))
logits = [5.0, 3.2, 2.9, 2.0, 1.5, -1.0]
draws = [sample(logits, temperature=2.0, top_p=0.9, rng=np.random.default_rng(s)) for s in range(1000)]
print(np.bincount(draws, minlength=6) / 1000) # Lyon (index 4) appears; banana never does
What the APIs actually accept
Parameter support differs between providers and, increasingly, between models from the same provider. The table reflects public documentation checked on 2 October 2026; treat it as a starting point and confirm against the current reference for the exact model you call.
| Interface | Temperature | Top-p | Top-k | Notes |
|---|---|---|---|---|
| Anthropic Messages API | 0 to 1, default 1 | yes | yes | Guidance is to adjust temperature or top_p, not both; some recent models reject requests that set both |
| OpenAI chat and responses | 0 to 2 | yes | not exposed | Reasoning models reject or ignore sampling parameters in many configurations |
| llama.cpp server | default 0.80 | default 0.95 | default 40 | Also min-p (default 0.05) and a configurable sampler order |
Two practical consequences follow. First, temperature 1.0 does not mean the same thing everywhere: on a 0-to-1 scale it is the maximum, on a 0-to-2 scale it is the middle, though in both cases it means the model's unmodified distribution. Second, reasoning models increasingly own their own sampling. If your client library always sends a temperature, it may break when you switch to one of them, so make sampling parameters optional per model in configuration rather than hard-coded.
Choosing settings by task
| Task | Starting point | Why |
|---|---|---|
| Classification, extraction, tool arguments | T = 0 or close to it | There is one right answer; variety is only error |
| Code generation, single answer | T = 0 to 0.3 | Low variance; raise only if you sample several candidates and test them |
| Several candidates for voting or tests | T = 0.6 to 1.0, top-p 0.95 | Candidates must differ to be worth generating |
| Conversational assistant | Provider default | Defaults are tuned for this; change only with evidence |
| Brainstorming, fiction | T around 1.0 with min-p or top-p | Variety is the product; truncation keeps it coherent |
Sampling is also not a substitute for structure. If output must parse as JSON, temperature 0 lowers the failure rate but does not guarantee validity; structured output and guided decoding mask invalid tokens so the guarantee holds at any temperature. If you want reliability from variety, self-consistency samples several reasoning paths at a moderate temperature and votes, which only works because the samples differ.
Temperature 0 is not reproducible
Setting temperature to 0 removes sampling randomness, but the logits themselves can change between identical requests. Floating-point addition is not associative, and a serving system that batches your request with different neighbours, or picks a different kernel for a different batch shape, can sum in a different order and shift a logit in the last bits. When two tokens are nearly tied, that flips the argmax, and every later token can differ. Mixture-of-experts routing can amplify the effect.
Some APIs offer a seed parameter, but providers describe it as best effort. If you need reproducibility, for an audit trail or a regression test, store the output rather than trying to regenerate it, and write tests that check properties of an answer rather than exact strings.
A tuning harness
Choose settings with a small evaluation rather than intuition. Fix the prompt and a set of 30 to 100 real inputs with a checker, then sweep a handful of settings and record two numbers: the pass rate, and how different the outputs are from each other.
import itertools, statistics
SETTINGS = [dict(temperature=t, top_p=p) for t, p in itertools.product([0.0, 0.3, 0.7, 1.0], [1.0, 0.9])]
def evaluate(call_model, cases, settings, samples=5):
rows = []
for s in settings:
passes, distinct = [], []
for case in cases:
outs = [call_model(case.prompt, **s) for _ in range(samples)]
passes += [case.check(o) for o in outs]
distinct.append(len(set(outs)) / samples)
rows.append((s, statistics.mean(passes), statistics.mean(distinct)))
return sorted(rows, key=lambda r: -r[1])
# Rule: take the lowest temperature whose pass rate is within noise of the best,
# unless the task needs variety, in which case require distinct >= your threshold.Five samples per case at temperature 0 also tell you how unstable the endpoint is: a distinct ratio above 0.2 there points at near-tied tokens, which usually means the prompt is ambiguous, not that the setting is wrong. The broader method for building checkers and comparing prompt variants is in prompt evaluation architecture.
Failure modes
- Repetition loops at low temperature: greedy decoding can fall into a cycle where repeating the last phrase is always the most likely continuation. Add a mild repetition penalty or raise the temperature slightly; do not add a stop sequence that hides the loop.
- Drift at high temperature: one unlikely token early can derail a long answer. Use min-p or top-p together with a high temperature, never temperature alone above 1.
- Settings copied across engines: different scales and sampler orders turn the same numbers into different behaviour. Re-run the harness after any model or provider change.
- Over-constrained stacks: top-k 10 with top-p 0.5 and T = 0.2 is greedy decoding with extra steps. Remove knobs until each one does something you can measure.
- Rejected parameters: sending temperature to a model that does not accept it fails the whole request. Validate the configuration per model at startup.
What to do next
- List every call site in your application and the sampling settings each sends, including defaults the client library adds.
- Classify each call site as single-answer, multi-candidate or creative, and set a starting point from the table above.
- Build a checker and 30 or more real inputs for your highest-volume call site, and run the harness across four to eight settings.
- Pick one truncation method per call site and remove the others unless the harness shows they help.
- Make sampling parameters optional per model in configuration, so a switch to a reasoning model does not send rejected fields.
- Replace exact-string regression tests with property checks, since temperature 0 is not reproducible.