What polysemanticity actually means
Polysemanticity is the property of a single unit — usually one neuron in an MLP hidden layer, indexed by one coordinate of the activation vector h ∈ R^d — responding strongly to several unrelated input features. Formally, if we write the neuron’s scalar activation as a_i = σ(w_i · x + b_i), a monosemantic neuron would have a_i large on exactly one coherent concept and near zero otherwise. A polysemantic neuron has a_i large on a grab-bag of concepts that share no semantic relationship.
The word borrows from linguistics, where a polysemous word like ‘bank’ carries multiple meanings. The analogy is loose but useful: just as context disambiguates ‘bank,’ the rest of the network’s activation vector supplies the context that tells downstream layers which of a neuron’s several meanings is currently in play. The neuron itself is genuinely overloaded — it is one wire carrying several signals that only the full pattern can separate.
A concrete picture
Make it tangible. Suppose you take a mid-network MLP neuron and collect the dataset examples that activate it most strongly — the standard ‘max activating examples’ view. For a monosemantic neuron you might see a clean list: every top example contains a period followed by a capital letter. Satisfying, readable, nameable.
For a polysemantic neuron you instead get something like: the top example is Python code indentation, the second is a Korean particle, the third is the token ‘theorem’ in math text, the fourth is a semicolon in a list. There is no honest single label. You could invent a tortured umbrella, but it would be a post-hoc story, not a mechanism. This is the everyday texture of interpreting raw neurons: a minority read cleanly, and the majority are these irreducible mixtures.
Superposition is the cause, polysemanticity is the symptom
The two terms are constantly conflated, and keeping them distinct is the whole conceptual payoff. Superposition is a claim about the representation: the model wants to represent more distinct features than it has dimensions, so it stores them as an overcomplete set of non-orthogonal directions in activation space, accepting a little mutual interference in exchange for capacity. Polysemanticity is the observable consequence you see when you look along the standard neuron basis.
The link is geometric. If m feature directions {f_1, …, f_m} are crammed into d < m dimensions, no single axis-aligned neuron can align with just one feature; each neuron ends up with non-trivial dot products f_k · e_i against many features at once. So the neuron lights up for all of them. Superposition is why the features overlap; polysemanticity is what that overlap looks like when you insist on reading one coordinate at a time. Cause and symptom.
Measuring it: how do you know a neuron is polysemantic?
Because there is no ground-truth list of ‘the features,’ measuring polysemanticity is empirical and triangulated. The workhorse is max-activating dataset examples: pull the inputs that maximize a_i and judge, by eye or with an automated labeler, whether they form one coherent concept or several. A neuron whose top examples split into distinct, unrelated clusters is polysemantic.
You can make that quantitative. Cluster the top-activating inputs in an embedding space and count well-separated clusters — roughly a count of distinct meanings per neuron. Another route is activation histograms and kurtosis: a clean monosemantic neuron often shows a sharply bimodal or heavy-tailed distribution tied to one trigger, while overloaded neurons blur. A third is causal: intervene on the neuron (ablate or clamp it) and check how many different downstream behaviors change. If knocking out one neuron degrades several unrelated capabilities, that neuron was carrying several features.
A small worked estimate
A back-of-envelope shows why overload is the default. Take an MLP hidden layer with d = 4096 neurons and suppose the model is genuinely tracking m = 40000 distinct interpretable features at this depth — a plausible order of magnitude for a capable model. Then even in the best case there are about m / d ≈ 10 features per neuron to account for.
features m = 40000
neurons d = 4096
features/neuron = m / d ≈ 9.8
if each feature is active on a fraction p = 0.01 of tokens,
expected simultaneous collisions per neuron ≈ (m/d) · p ≈ 0.1The last line is the trick: with ~10 features sharing a neuron but each firing only 1% of the time, the expected number that fire together on any token is about 0.1 — usually zero, occasionally one. The neuron can host ten meanings and still present a clean, single-feature signal on almost every token. Sparsity is what makes the overload nearly free.
Why polysemanticity obstructs interpretability
Mechanistic interpretability wants to name the parts of a network and compose those names into an explanation of behavior — the way you read a program by reading its variables and functions. Polysemanticity breaks the first step. If a neuron has no single meaning, there is no honest label to write on it, and every explanation built from neuron labels inherits the ambiguity.
It also poisons interventions. A favourite tool is ablation: zero out a neuron and see what breaks, inferring that the neuron was responsible for whatever changed. But ablating a polysemantic neuron perturbs every feature it participates in, so the observed effect is a superposition of side effects you cannot cleanly attribute. Probing suffers the mirror problem: a linear probe can read a concept off a population of neurons even though no individual neuron represents it, tempting you to credit the wrong units. The neuron, it turns out, is simply the wrong basis in which to look.