The question a fixed pipeline never asks

A standard RAG system retrieves unconditionally, embedding an assumption: that the model does not know the answer and that the corpus does. Both are often false. For a question the model would have answered correctly from parameters, retrieval adds latency, tokens, and a real chance of harm — an irrelevant passage can pull a correct answer off course. For a multi-hop question, one round is not enough: the second hop depends on a fact you only learn from the first, so a single shot at the index cannot surface the bridging document.

A fixed k-chunk, one-round pipeline is therefore simultaneously too much machinery for the easy queries and too little for the hard ones. Adaptive RAG’s premise is that query difficulty is predictable in advance — cheaply, from the query text alone, before you pay for any retrieval or generation. If that prediction is even moderately accurate, route each query to the cheapest strategy that will actually answer it.

Advertisement

Three routes, one classifier

The canonical formulation (Jeong et al., Adaptive-RAG, NAACL 2024) defines three strategies of increasing cost:

RouteWhat it doesFits
A — no retrievalAnswer straight from parametersCommon knowledge, definitions, arithmetic, rewriting
B — single-stepOne retrieval, one generationOne fact lives in one document
C — multi-stepIterative retrieve→reason→retrieveMulti-hop, comparative, aggregative questions

A small classifier f(q) → {A, B, C} sees only the query string and picks one. Note what this is not: it is not a reranker (it never sees documents) and not a self-critique loop (it runs before any generation). It is a pure feed-forward routing decision, which is exactly why it can be small and fast. The two fixed baselines are degenerate cases of the same design — routers that always output B, or always output C.

Advertisement

The expected-cost model

The whole argument is arithmetic. Let c_A, c_B, c_C be the end-to-end cost of each route (measured however you like: milliseconds, generated tokens, dollars), and let p_A, p_B, p_C be the fraction of traffic the router sends down each. Ignoring the router itself for a moment:

E[cost] = Σ_r p_r · c_r  = p_A·c_A + p_B·c_B + p_C·c_C

always-single-step:  cost = c_B
always-multi-step:   cost = c_C
adaptive:            cost = c_router + Σ_r p_r · c_r

The router costs c_router on every query, so it must be cheap relative to the spread between routes. Since c_A < c_B < c_C and typically c_C ≈ 3–5 × c_B (each hop is another retrieval plus another generation), the savings come from two places: queries diverted from B to A, and queries kept out of C.

A worked example: 1000 queries

Put numbers on a mixed workload. Say single-step costs c_B = 900 ms, no-retrieval costs c_A = 400 ms (shorter prompt, generation only), and a three-hop route costs c_C = 2900 ms. A DistilBERT-class router adds c_router = 12 ms. Traffic splits p_A = 0.35, p_B = 0.50, p_C = 0.15.

adaptive  = 12 + 0.35(400) + 0.50(900) + 0.15(2900)
          = 12 + 140 + 450 + 435  = 1037 ms

always-B  = 900 ms      (but fails the 15% multi-hop tail)
always-C  = 2900 ms     (correct everywhere, 2.8× the cost)

Read that carefully: the honest conclusion is not “adaptive is cheapest.” Adaptive is 15% slower than always-single-step, and 2.8× cheaper than always-multi-step while matching its accuracy on the hard tail. That is the real trade — multi-step-quality answers at close to single-step price. The value scales with how heterogeneous your traffic is; on a workload that is 95% simple lookups, a router is overhead with nothing to save.

Where the labels come from

Nobody hand-labels queries as easy or hard, and the labels you want are not about the query’s surface form anyway — they are about your model on your corpus. Derive them from outcomes instead: run every training query through all three routes and record which ones answer it correctly.

A correct?  → label A   (cheapest sufficient route)
else B correct?  → label B
else C correct?  → label C
none correct     → fall back to the dataset prior
                   (single-hop set → B, multi-hop set → C)

This is silver labelling: cheap, automatic, and self-calibrating, because a stronger generator naturally produces more A labels and shifts the distribution toward the cheap route. The cost is one offline sweep at full multi-step price. What remains is a plain three-way text classifier fine-tuned on (query, label) pairs.

Routing errors are not symmetric

A router that is 85% accurate is not 15% bad, because the two error directions have different consequences. Under-routing (predicting A or B for a query that needed C) yields a confidently wrong answer: you saved 2 seconds and returned bad information. Over-routing (predicting C where A sufficed) yields a correct answer that cost too much. One error damages the product; the other damages the bill.

Because of that asymmetry you should almost never deploy the raw argmax. Bias toward the expensive route with a class-weighted threshold — take route A only when P(A | q) > τ with τ ≈ 0.7, not merely when A is the top class. This trades some cost saving for fewer confidently-wrong answers. Tune τ by plotting accuracy against expected cost on a held-out set and picking the knee — not by maximizing classifier accuracy, which is a lossy proxy.

Confidence-triggered retrieval: the other adaptive axis

Query-side routing decides before generating. A second family decides during generation, using the model’s own uncertainty as the trigger. FLARE-style methods draft a sentence speculatively, inspect its token probabilities, and retrieve only if confidence dips: if min_t p(x_t) < θ over the drafted span, discard it, retrieve using the draft as the query, and regenerate.

The appeal is that it needs no trained router and adapts within a single long answer — paragraph three may need a citation even if paragraphs one and two did not. The cost is that you pay for speculative tokens you sometimes throw away, and that raw token probability is a mediocre proxy for factual uncertainty — models are routinely fluent and confident while wrong. In practice the two axes compose: a query-level router picks the strategy, and a confidence trigger inside the long-form routes decides when to reach for the index again.

What actually costs what

Routing decisions only make sense against a real cost profile, and the profile is usually more lopsided than people expect. For a query of N_q tokens, k retrieved chunks of L tokens each, and an answer of N_out tokens:

embed query      ~ O(N_q · d)          — microseconds
ANN search       ~ O(log M · d)       — ~1–10 ms over M vectors
prefill context  ~ O((N_q + k·L)·d^2)  — grows with k·L
decode answer    ~ O(N_out · d^2)      — memory-bandwidth bound

multi-step: multiply the whole stack by H hops

Vector search is almost never the bottleneck. The dominant terms are prefill over the retrieved context and decoding, both driven by how many tokens you put in front of the model. That is why route A is so much cheaper than route B: not because it skips the index, but because it drops k·L context tokens — typically 2000–4000 of them — from prefill entirely.