Decode is memory-bound — why draft-then-verify wins
Autoregressive decode generates one token per forward pass, and each pass must stream the entire weight matrix from memory to compute a single new position. The arithmetic is trivial; the memory traffic is not. So decode is memory-bandwidth-bound: the hardware’s compute units sit mostly idle, waiting on weights. A forward pass that scores one position and one that scores a dozen cost almost the same wall-clock time, because both are gated by loading the weights, not by the flops.
Speculative decoding exploits exactly this slack. Cheaply guess the next γ tokens with some fast drafter, then run the expensive target model once to score all γ+1 candidate positions in parallel. Verification is lossless: a rejection-sampling check accepts the longest correct prefix and preserves the target’s exact output distribution. If the guesses are good you collect several tokens for the price of one weight-load. The whole game is a drafter that is both cheap and accurate.
Vanilla speculative decoding: a separate draft model
The original recipe uses a second, smaller model q to draft tokens autoregressively, then the target p verifies them in one pass. Each drafted token x is accepted with probability min(1, p(x)/q(x)); on the first rejection you resample that position from the adjusted distribution (p − q)_+ and stop. This is provably equivalent to sampling from p directly — no quality loss.
Let α be the average per-token acceptance probability. The expected number of tokens produced per verification pass is a geometric sum:
E = (1 − α^(γ+1)) / (1 − α)The +1 is the free bonus token the verification pass always yields. The catch is practical: you must find or train a separate draft model whose distribution is close enough to the target to earn a high α, keep it aligned across fine-tunes, and pay to deploy two models. A good small drafter for a given target is genuinely hard to come by.
Medusa: independent token heads
Medusa removes the second model. It freezes the target and attaches K extra lightweight heads on top of the last hidden state h_t; head k predicts the token at position t+k directly, and all heads fire in parallel from the same h_t. No autoregression, no separate network — one forward pass emits several draft tokens at once.
The weakness is structural. Head 2 predicts the token two steps out without knowing what head 1 actually produced. It is modelling the marginal p(t+2) rather than the conditional p(t+2 | t+1), yet real text is highly conditional — the next word depends on the word just chosen. Because each head ignores its predecessors’ realized outputs, joint accuracy decays quickly with distance and acceptance falls off. Medusa recovers some of this by verifying a tree of candidate combinations, but the root cause — independent, non-autoregressive heads — remains. That is the gap EAGLE closes.
EAGLE’s move: draft at the feature level
EAGLE’s insight is that tokens are the wrong thing to draft. A token is a discrete, high-entropy sample; the vector that produces it is smoother and more regular. That vector is the feature f — the second-to-top hidden state, the output of the final decoder layer that the LM head consumes to form the next-token distribution. Feature sequences evolve far more predictably than token sequences, so extrapolating one step ahead in feature space is an easier learning problem than guessing the next token outright.
So EAGLE runs a single small autoregression head — one extra decoder layer — over the sequence of features, predicting the next feature f′ from the current one. It then applies the target’s own, unchanged LM head to f′ to get a token distribution and samples from it. No second model, no independent heads: just cheap autoregression in feature space, decoded through the very head the target already uses. Reusing that exact LM head is a big part of why the drafts stay close to the target’s distribution.
The uncertainty problem, and the token that fixes it
There is a subtlety that would sink naive feature extrapolation. The map from one feature to the next is not deterministic. A feature f yields a distribution over next tokens; which token is actually sampled is part of the state, and it changes the next feature. So f alone underdetermines f′ — the same feature can lead to different continuations depending on the draw. This is the exact ambiguity Medusa’s independent heads swallow.
EAGLE resolves it by feeding the head the token that was actually sampled. One draft step holds the pair (f, t) and advances it:
f′ = AR_Head( f , e(t) ) # e(t): embedding of the sampled token
p = softmax( LM_head(f′) ) # target’s own head, unchanged
t′ ~ p # sample the next token
(f, t) ← (f′, t′) # advance one step, repeat γ timesThe head’s input is the feature concatenated with the shifted token embedding, so it always knows which branch of the sampling was taken. That single conditioning term collapses the uncertainty that non-autoregressive heads cannot.
Why feature drafting lands a higher acceptance rate
Two things push EAGLE’s per-token distribution close to the target’s, which is what raises α. First, it decodes drafted features through the target’s exact LM head, so the final feature→token mapping is identical rather than approximated by a different network. Second, it is genuinely autoregressive and conditions on the realized token, so it models p(t+2 | t+1) instead of the marginal. A separate draft model suffers a distribution mismatch; Medusa’s heads suffer a dependency mismatch; EAGLE has neither.
The head is trained cheaply on the target’s own recorded features with a two-part loss: a smooth-L1 regression term that makes the predicted feature f′ track the true next feature, plus a cross-entropy term on the decoded token distribution. Aligning both the vector and the distribution lets one small layer imitate the target well enough to earn acceptance rates well above token-level or independent-head drafting — and, as the next section shows, acceptance rate is where all the speedup lives.