Features are directions, and there are too many of them
The starting assumption is the linear representation hypothesis: a feature — some human-meaningful property like ‘this text is legal boilerplate’ or ‘the subject is plural’ — is represented as a direction W_i ∈ R^d, and the activation vector is a weighted sum of the directions currently present:
x = Σ_i a_i · W_i x: [d], W_i: [d], ||W_i|| = 1If the W_i were mutually orthogonal you could read a_i back exactly with a dot product, and the model would be limited to d features — 768 concepts for a 768-wide stream. That limit is plainly violated. Probing work recovers thousands of distinct, causally-load-bearing directions from streams narrower than the feature count, and dictionary-learning runs surface tens of thousands. So either the linear hypothesis is wrong, or the directions are not orthogonal. Superposition takes the second branch: keep linearity, give up orthogonality, and treat the cross-talk as a cost to be budgeted.
The toy model: sparse features through a bottleneck
The cleanest evidence comes from a model small enough to fully understand. Take m synthetic features, project them down to n < m dimensions, and ask the network to reconstruct them:
x ∈ R^m x_i = 0 with prob. S, else Uniform(0,1)
h = W x W: [n, m] (the bottleneck)
x̂ = ReLU( W^T h + b ) = ReLU( W^T W x + b )
L = Σ_i I_i · (x_i - x̂_i)^2 I_i = per-feature importanceEverything interesting lives in the m × m Gram matrix W^T W. It cannot be the identity — its rank is at most n — so off-diagonal entries are forced to be nonzero. The two dials are sparsity S and importance I_i. At S = 0 (features always on) the model learns the top-n principal directions and discards the rest, exactly as PCA would. Crank S toward 1 and it starts representing all m features anyway. That switch, in a model with three lines of algebra, is superposition.
How many near-orthogonal directions fit in d dimensions?
The capacity argument is Johnson–Lindenstrauss in spirit. Two independent random unit vectors in R^d have an inner product of mean 0 and standard deviation 1/√d, with a sub-Gaussian tail:
P( |<u, v>| > ε ) ≤ 2 exp( -d ε^2 / 2 )
union bound over the m(m-1)/2 pairs → all pairs stay below ε if
m < exp( d ε^2 / 4 )The count is exponential in d, not linear — and quadratic in the interference you tolerate. For d = 768:
| Tolerated cosine ε | Guaranteed directions m |
|---|---|
| 0.10 | ≈ 7 |
| 0.20 | ≈ 2.2 × 10^3 |
| 0.30 | ≈ 3.2 × 10^7 |
| 0.40 | ≈ 2.2 × 10^13 |
The union bound is loose, so read these as sufficient conditions, not ceilings. The shape is the point: nudging the tolerable cosine from 0.2 to 0.3 buys four orders of magnitude. Geometry is not the binding constraint here.
Interference: what non-orthogonality actually costs
Read feature i back with its own direction and the error term falls straight out of the Gram matrix:
(W^T W x)_i = ||W_i||^2 x_i + Σ_(j≠i) (W_i · W_j) x_j
\___________/ \_______________________/
signal interferenceThe interference is a sum of k independent-ish terms, where k is the number of co-active features. With random directions the cosines have scale 1/√d, so the noise grows as √k / √d — it depends on how many features fire together, not on how many exist.
Worked example: d = 768, k = 50 co-active features of typical magnitude 0.5. Noise standard deviation ≈ 0.5 · √50 / √768 ≈ 0.13. That is why the decoder’s ReLU(· + b) matters: a learned negative bias around -0.13 clips the noise on inactive features to exactly zero, at the cost of shaving the same amount off every active one. Superposition is only survivable because a threshold nonlinearity is there to absorb it.
The sparsity threshold: when the trade turns favourable
Take the smallest interesting case: two features, one dimension. The dedicated strategy stores feature 1 and drops feature 2, paying full error whenever feature 2 fires. The superposed strategy is antipodal — W_1 = +1, W_2 = -1 — and recovers either feature perfectly when it fires alone, failing only on a collision.
p = P(a feature is active) = 1 - S
dedicated loss ≈ p · E[x^2] (feature 2 always lost)
superposed loss ≈ c p^2 · E[x^2] (both on: prob p^2)
superposition wins when c · p < 1Collisions are quadratic in p while dropped features are linear, so sparsity is exactly the lever. Scaled up, the criterion is that the expected number of simultaneously active features stays under the dimension count: m · p < d. Concretely, m = 100,000 features each firing on 0.1% of tokens gives m p = 100 against d = 768 — comfortably inside budget, at 130× more features than dimensions. Superposition is not exotic; at realistic sparsity it is the obvious move.
Phase transitions and the geometry that emerges
Sweep sparsity in the toy model and features do not fade smoothly into shared directions — they snap. Define a feature’s effective dimensionality as the share of a dimension it owns:
D_i = ||W_i||^2 / Σ_j ( Ŵ_i · W_j )^2 Σ_i D_i ≤ nD_i lands on a small set of sticky rational values, each corresponding to a uniform polytope: 1 for a dedicated dimension, 1/2 for an antipodal pair sharing one axis, 2/3 for three features at 120° in a plane, 3/4 for a tetrahedron in three dimensions, 2/5 for a pentagon in a plane. Between them are gaps — genuine phase transitions, with sudden reorganisation as sparsity or importance crosses a boundary.
The constraint Σ_i D_i ≤ n is the real content: capacity is a conserved budget of n dimension-shares, allocated by importance. Important features buy exclusive axes; the rest crowd into polytopes; the least important get nothing at all.
Polysemantic neurons are the observable consequence
Superposition is a claim about geometry, but what you actually see when you probe a model is polysemanticity. Read the value of neuron j and you get a projection onto a basis vector, not onto a feature:
x_j = Σ_i a_i · (W_i)_jEvery feature whose direction has a nonzero j-th component contributes. When m >> d, the directions cannot align with the axes, so a single neuron is a weighted mixture of many unrelated features and its activation profile looks like nonsense: high on Python decorators, high on chemical formulae, high on the word ‘however’.
This reframes a long-standing observation. Polysemantic neurons were treated as evidence that networks are inscrutable. Superposition says the opposite: the representation is clean and linear, you are reading it in the wrong basis. Interpretability becomes a change-of-basis problem — recover W, and monosemantic units reappear.