Reflection tokens: critique as vocabulary
The central trick is to stop treating ‘should I retrieve?’ as an external policy and make it a token. Self-RAG extends the vocabulary V with a small set of reflection symbols V_r, so the model decodes over V ∪ V_r with one unchanged softmax:
[Retrieve] ∈ {yes, no, continue}
[IsRel] ∈ {relevant, irrelevant} ← per retrieved passage
[IsSup] ∈ {full, partial, none} ← per generated segment
[IsUse] ∈ {1, 2, 3, 4, 5} ← per generated segment
p(r | x, y_<t) = exp(z_r) / Σ_{v ∈ V ∪ V_r} exp(z_v)Nothing about the architecture changes — W_out simply grows from [d, |V|] to [d, |V| + |V_r|], a dozen extra rows. The payoff is that a critique is now a probability, not a second model’s opinion. That single design choice is what makes every knob later in this article a scalar you can turn.
The retrieve-or-not decision
At the start of each segment the model emits [Retrieve]. Normalising over just the yes/no branch gives a calibrated urge to search, which you compare against a threshold δ:
p_ret = p(yes) / ( p(yes) + p(no) )
retrieve if p_ret > δ (δ ≈ 0.2 retrieves often; δ ≈ 0.8 rarely)Why not always retrieve? Because retrieval has a real cost function. Writing a for accuracy, adding k passages costs k · L_c prefill tokens and injects distractors; the expected gain is positive only when the parametric knowledge is weak. For ‘who wrote Hamlet’ the model already knows, and a passage about a Hamlet, Ontario town hall is pure noise. Empirically the loss is real: on closed-book questions a model already answers correctly, forced retrieval flips a meaningful slice of them wrong. [Retrieve] exists to skip those.
Per-passage relevance critique
When retrieval fires, the retriever returns K passages d_1 … d_K. Self-RAG does not concatenate them. Each passage opens its own parallel continuation, and the model emits [IsRel] conditioned on that passage alone:
s_rel(d_i) = p(relevant | x, d_i) / [ p(relevant | x, d_i) + p(irrelevant | x, d_i) ]This is a cross-encoder-style judgement — query and passage share the same attention pass — but produced by the generator itself, so it is free of the extra reranker model. The structural win is isolation: in concatenated RAG one irrelevant passage can dominate attention and poison the answer, because the model cannot mark it off. Here each branch is scored independently, so a bad passage costs you one dead branch instead of a corrupted context. A passage with s_rel below threshold is pruned before it ever produces text. The cost of that isolation is that the branches no longer see each other, so a fact that only emerges from combining two passages is lost — genuinely multi-hop questions still need an iterative loop on top.
Groundedness: scoring the answer against the evidence
[IsSup] is the token that makes Self-RAG more than routing. After a segment y_t is drafted from passage d_i, the model grades the entailment d_i → y_t on three levels: fully supported (every claim traceable to d_i), partially supported (some claims are parametric), no support (contradicted or invented). It is scored as a weighted expectation:
s_sup = 1.0·p(full) + 0.5·p(partial) + 0.0·p(none)Note the direction of the check. [IsRel] asks whether the evidence suits the question; [IsSup] asks whether the answer is licensed by the evidence. Those come apart constantly: a perfectly relevant passage can accompany a fluent, confidently hallucinated sentence. [IsSup] is the only signal in the pipeline that looks at the generated text — every retrieval-side improvement in the world leaves hallucination untouched, because hallucination happens downstream of retrieval. The partial level is the interesting one: it is the honest label for a sentence that blends a retrieved fact with a parametric one, which is what most useful answers actually are.
Utility, and why it is a separate token
A sentence can be flawlessly grounded and still useless — quoting a definition when the user asked for a comparison, or answering a narrower question than the one posed. [IsUse] grades that on a 1–5 scale, collapsed to a scalar by expectation over the five token probabilities:
s_use = Σ_{i=1..5} w_i · p(i), w = (−1, −0.5, 0, +0.5, +1)Keeping utility separate from support is deliberate, and it is the axis most RAG evaluations conflate. Optimising groundedness alone produces a timid system that paraphrases whichever passage it holds and never commits; optimising utility alone produces a confident one that drifts off the evidence. The two scores pull in opposite directions by design, and the next section is where you choose the exchange rate between them. Treating them as one number — a single ‘quality’ score — hides exactly the tension you most need to control.
The segment-level beam score, worked
Self-RAG decodes a segment at a time. For each candidate continuation — one per surviving passage, plus the no-retrieval branch — the rank score adds the language-model term to a weighted sum of critiques:
S(y_t, d_i) = log p(y_t | x, d_i) + w_rel·s_rel + w_sup·s_sup + w_use·s_useTake weights (1.0, 1.0, 0.5) and two branches. Branch A: log p = −0.90, s_rel = 0.95, s_sup = 0.40, s_use = 0.80 → −0.90 + 0.95 + 0.40 + 0.40 = 0.85. Branch B: log p = −1.30, s_rel = 0.80, s_sup = 0.95, s_use = 0.60 → −1.30 + 0.80 + 0.95 + 0.30 = 0.75. The fluent-but-unsupported branch A wins — until you raise w_sup to 2.0, which lifts B to 1.70 against A’s 1.25. The tradeoff is a dial, not a property of the model.