What actually limits the context window
Nothing in a transformer’s parameter shapes forbids long sequences. W_Q, W_K, W_V are [d, d] regardless of N; attention softmax(QK^T / sqrt(d_k)) V is defined for any N. The wall is statistical, not structural.
With RoPE, position enters as a rotation angle m · θ_i for token index m and frequency θ_i. During pretraining at length L, the model only ever observes angles in [0, L·θ_i] for each i. Evaluate at m > L and the low-frequency dimensions produce phases that never appeared in any training batch. The dot product between a query and a key is a sum of cosines of relative phase; feed it unseen phases and it returns values the network has no calibration for. Perplexity does not drift upward — it explodes, often within a few hundred tokens past L.
RoPE's frequency ladder, with numbers
RoPE splits a head of dimension d into d/2 pairs, each rotated at its own rate:
θ_i = b^(-2i/d), i = 0 .. d/2 - 1, b = 10000 (typical)
wavelength λ_i = 2π / θ_i
rotations over training context: r_i = L / λ_iTake d = 128, b = 10000, L = 4096. The top pair, i = 0, has θ_0 = 1, so λ_0 ≈ 6.28 tokens and r_0 ≈ 652 full rotations — it has seen every phase thousands of times over. The bottom pair, i = 63, has θ_63 = 10000^(-0.984) ≈ 1.16×10^-4, giving λ_63 ≈ 54,400 tokens and r_63 ≈ 0.075 — it completes less than a tenth of one turn across the entire training window. That single number is the whole problem.
The out-of-distribution phase, quantified
Follow dimension 63. Its largest angle ever seen in training is 4096 × 1.16×10^-4 ≈ 0.47 rad, about 27°. Ask the model to run to 32k and position 32768 demands 32768 × 1.16×10^-4 ≈ 3.79 rad, about 217° — more than half a turn, in a region of the circle the pair has literally never occupied.
The high-frequency pairs are fine: they wrapped hundreds of times, so any new angle is one they have seen. It is the slow dimensions — the ones encoding coarse, long-range position — that break, and they break precisely because they were designed not to wrap. So the design problem is sharp and narrow: keep the fast dimensions untouched, and do something about the slow ones before they leave the arc they know. Every method below is a different answer to that one sentence.
Position Interpolation: squeeze instead of extrapolate
The first working fix (Chen et al., 2023) is almost embarrassingly simple. To reach length L’ = s·L, do not let positions run past L — rescale them into the old range:
m → m / s equivalently θ_i → θ_i / s
s = L’ / L (extension factor, e.g. 4096 → 32768 gives s = 8)Now position 32768 maps to angle 4096·θ_i — in distribution by construction. Interpolation is far safer than extrapolation, and PI works with roughly 1000 fine-tuning steps.
The cost is paid at the top of the ladder. Adjacent tokens used to be separated by θ_0 = 1 rad; now they differ by 1/8 rad. The model’s ability to tell ‘previous token’ from ‘three tokens back’ is compressed by the same factor s that bought the range. That is why plain PI degrades short-context quality, and why every later method is essentially a way to stop squeezing the fast dimensions.
NTK-aware scaling: move the base, not the positions
The NTK-aware trick changes a single scalar — the base b — so that interpolation is applied unevenly across the ladder:
b’ = b · s^(d / (d - 2))
top pair (i = 0): θ’_0 = b’^0 = 1 → unchanged
bottom pair (i = d/2 - 1): θ’_i = θ_i / s → fully interpolatedThe exponent is chosen to make that second line exact. At i = d/2 - 1 the exponent 2i/d = (d-2)/d, so b’^((d-2)/d) = b^((d-2)/d) · s — the s^(d/(d-2)) and the (d-2)/d cancel to leave exactly one factor of s. For d = 128, s = 8: b’ = 10000 · 8^(64/63) ≈ 82,700. Dimensions in between get a smooth, monotonic blend. One number changed, and the local resolution loss that hurt PI largely disappears — it even works passably without any fine-tuning.
Dynamic NTK and the brute-force base bump
Two practical variants deserve naming. Dynamic NTK makes the scale a function of the current sequence: s = max(1, N_current / L), recomputed as generation proceeds. Below the original length the model is bit-identical to the one that was pretrained, so there is zero short-context regression; the correction only ramps in when it is needed.
Adjusted base frequency (ABF) skips the cleverness and simply pretrains or continues training with a much larger base — Code Llama used b = 10^6, Llama 3 uses b = 5×10^5. A larger base stretches every wavelength, so the slow dimensions wrap even less and their phases stay well inside the trained arc at long N. It costs real training compute rather than a post-hoc edit, but when you control pretraining it is the sturdiest option, and it composes with the methods below.
YaRN: interpolate by parts
YaRN (Peng et al., 2023) makes the ‘fast dimensions untouched, slow dimensions interpolated’ intuition explicit, using rotations-per-context r_i = L/λ_i as the classifier:
γ_i = clamp( (r_i - α) / (β - α), 0, 1 ) typical α = 1, β = 32
θ’_i = (1 - γ_i) · (θ_i / s) + γ_i · θ_iRead it as three regimes. r_i > β (many rotations, high frequency): γ = 1, leave alone. r_i < α (under one rotation, the dangerous ones): γ = 0, interpolate fully. Between them, a linear ramp. With d = 128, b = 10000, L = 4096, the cutoffs land at i ≈ 21 and i ≈ 45 — roughly the fastest third untouched, the slowest third fully squeezed, and a ramp across the middle.
YaRN&amp;amp;#x27;s second half: attention temperature
The frequency fix is only part of YaRN. Lengthening the context also changes the softmax itself: a query now scores against s× more keys, and the entropy of the resulting distribution grows roughly like ln N. Attention that was crisply peaked at 4k becomes diffuse at 32k — a diluted average over far more candidates.
YaRN counters it by sharpening the logits with a scale fitted empirically across extension factors:
attn = softmax( (q · k) / (t · sqrt(d_k)) ), sqrt(1/t) = 0.1 · ln(s) + 1
s = 8 → sqrt(1/t) = 0.1(2.079) + 1 = 1.208
q and k each scaled by 1.208, so the logits scale by 1/t ≈ 1.46The elegance is that a constant multiplier on q and k can be folded straight into the precomputed cos/sin tables, so the correction costs nothing at inference and needs no code change in the attention kernel.