What a probe actually is
Formally, fix a trained transformer and pick a layer ℓ. For each input you extract a hidden vector h ∈ R^d — a token representation, or a pooled sentence vector. You do not touch the model’s weights. Instead you train a separate classifier gθ: R^d → Y on labelled pairs (h, y), where y is the property you care about.
A linear probe is just g(h) = softmax(W h + b) with W: [|Y|, d], trained by cross-entropy on a held-out split of activations, and you report its test accuracy. The logic is simple: the frozen h is the only thing the probe sees, so any signal it exploits must already be sitting in h. A linear probe additionally constrains how that signal is stored — if a plain linear map recovers y, the property lies along a readable direction in activation space, not buried in a tangled nonlinear fold.
Why linear, specifically
The restriction to a linear map is the whole point, not a convenience. The quantity people care about in mechanistic interpretability is linear decodability, because transformer components read from the residual stream through linear projections. An attention head’s query and a next layer’s weight matrix both take linear functions of h. So a feature the model can act on in one step is, almost by definition, one a linear probe can read.
This connects to the linear representation hypothesis: that high-level concepts are stored as directions, so (v · h) / ||v|| along a concept vector v tracks how much of that concept is active. A linear probe is exactly a supervised way to find such a v. If you instead need a two-layer MLP to hit good accuracy, the information is present but nonlinearly entangled — a weaker claim.
A worked example
Suppose a small model tracks whether a running counter is odd or even, and you probe the residual stream at layer 6 with d = 256. You collect 20,000 token activations, split 80/20, and fit a logistic-regression probe. It reaches 98% test accuracy.
Read this carefully. The claim you have earned is narrow: parity is linearly decodable at layer 6. You have not shown the model computes its output using this direction, nor that parity is stored only here, nor that the 2% errors are noise rather than a systematic blind spot. You have also not ruled out that the probe latched onto a correlate — maybe token position correlates with parity in your data, and the probe read position instead. The single number 0.98 compresses all of these distinct possibilities into one digit, which is why probing needs the controls that follow.
Probe accuracy is not causal use
The central caveat: a property being decodable does not mean the model uses it. A representation can carry information as a harmless byproduct — a residual of the input distribution, a spandrel of training — that never feeds a downstream computation. Your probe finds it; the model ignores it.
The distinction is correlational vs causal. Probing is a correlational method: it measures mutual information between activations and a label, filtered through the probe’s hypothesis class. To claim the model relies on a feature you need an intervention — edit the activation and see whether behaviour changes. This is why probing pairs naturally with activation patching and causal analysis: the probe locates where information plausibly lives and points the intervention, and the intervention tests whether it is load-bearing. Treating a high probe accuracy as proof of mechanism is the most common error in the literature, and the one every subsequent control is designed to guard against.
The expressivity confound
If a linear probe is good, why not use a powerful MLP probe and decode even more? Because probe capacity is a confound. A sufficiently expressive probe can learn the task itself from the activations rather than reading a property the model already computed. Push far enough and a probe trained on random, untrained-network activations still scores well — it memorised the mapping using the activations as arbitrary features.
So high accuracy from a strong probe is ambiguous: it could mean the representation encodes the property, or merely that the probe is a good learner. This is the expressivity–faithfulness tradeoff. The more capacity you give the probe to extract information, the less its success tells you about the representation, because the credit shifts from the model to the probe. Keeping probes linear is one defence; the sharper defence is to measure the probe against a baseline that isolates how much work the probe is doing on its own.
Control tasks and selectivity
Hewitt and Liang (2019) formalised that defence. Alongside your real task, define a control task: the same inputs, but with labels assigned randomly and fixed per word type. The control has no linguistic structure, so any accuracy a probe gets on it comes purely from the probe’s own capacity to memorise, not from the representation.
Their key metric is selectivity: selectivity = acc(real task) − acc(control task). A probe that scores 97% on the real task but 89% on the random control has selectivity 8 — most of its success is memorisation, and it tells you little about the model. A probe that scores 90% real and 20% control has selectivity 70: it genuinely reads structure the representation supplies. The lesson is that raw accuracy is uninterpretable in isolation. A trustworthy probe is one that is simple enough to fail the control while still succeeding on the real task — which pushes you toward linear or tightly regularised probes.