The setup: one layer, one modular sum
The whole story lives in a tiny system. Take a one-layer transformer and train it on a single task: given two tokens a and b drawn from {0, …, P-1} with P = 113 (a prime), predict c = (a + b) mod P. The input is the sequence a b = and the model outputs the right residue over 113 logits, with a fixed fraction of the 113×113 pairs held out as a test set.
The task is chosen because it is small enough to fully understand yet rich enough to grok. Modular addition has a clean group structure, a known correct algorithm, and a network with enough capacity can trivially memorize the training pairs. So the question is sharp: when the model finally generalizes to unseen (a, b) pairs, what did it actually compute? Because the model is one layer with a handful of heads and one MLP, that question has an answer you can write down.
The answer in one line: add by multiplying waves
The reverse-engineered algorithm is a three-step trick built on trigonometry. Instead of treating a and b as integers, the model turns each into sinusoids at a few special key frequencies w_k = 2πk/P. It then multiplies those sinusoids, and multiplying waves of the same frequency is exactly the operation that adds their angles, via the identity
cos(w(a+b)) = cos(wa)cos(wb) − sin(wa)sin(wb)
sin(w(a+b)) = sin(wa)cos(wb) + cos(wa)sin(wb)So the network never ‘knows’ that a + b equals some integer. It represents a and b as angles on a circle, multiplies to get the angle w(a+b), and reads off which residue c that angle points to. Modular arithmetic falls out for free, because angles are periodic — going once around the circle is exactly ‘mod P.’ The rest of this article expands each step and shows the evidence that this, not memorization, is what the grokked model runs.
Step 1: embeddings are sparse in the Fourier basis
The first surprise is in the embedding matrix. Each token a has a learned embedding vector; stack them into a [P, d] matrix and take the discrete Fourier transform down the token axis. In a memorizing network this spectrum is dense — every frequency carries some weight. In the grokked network it is strikingly sparse: almost all the energy sits at a small handful of frequencies (roughly five or six of the P possible ones). Those are the key frequencies w_k.
Concretely, each embedding direction behaves like cos(w_k a) or sin(w_k a) for one of the key frequencies. The model has, on its own, chosen to represent an integer a not as a magnitude but as its coordinates on a few circles of different winding rates: (cos w_k a, sin w_k a). This is a genuine change of basis learned by gradient descent, and the first quantitative fingerprint of the generalizing circuit — the sparsity of that spectrum can be measured and watched over training, which is exactly what the progress measures below do.
Step 2: attention and the MLP multiply the waves
Having a and b encoded as sinusoids, the middle of the network multiplies them. Products of the input coordinates — terms like cos(w_k a)·cos(w_k b) and sin(w_k a)·sin(w_k b) — arise because attention mixes the a and b positions into the = position and the MLP’s nonlinearity produces the cross terms. In combination the MLP neurons compute cos(w_k(a+b)) and sin(w_k(a+b)) for each key frequency, precisely by the angle-addition identities above.
You can see this directly: individual MLP neurons, plotted as a function of (a, b), are periodic — they light up in a regular striped pattern whose period matches one of the key frequencies. A neuron is not detecting ‘a = 37’; it responds to a + b at frequency w_k. The activations are a smooth trigonometric surface, not a lookup table of memorized pairs, which is why the circuit extends gracefully to pairs it never saw in training.
Step 3: the readout sums cosines to a peak
The final step turns the summed angle w_k(a+b) back into a residue. The unembedding correlates the internal cos(w_k(a+b)) and sin(w_k(a+b)) against cos(w_k c) and sin(w_k c) for each candidate output c. Using the cosine angle-subtraction identity, that inner product collapses to a single clean expression for the logit of class c:
logit(c) = Σ_k cos(w_k · (a + b − c))Read this carefully, because it is the punchline. Every cosine term is maximized (equal to 1) exactly when a + b − c ≡ 0 (mod P), i.e. when c = (a+b) mod P. For any other c the frequencies point in different directions and partly cancel. Summing over several key frequencies sharpens the peak and suppresses off-target classes — they act like an error-correcting code, agreeing only at the true answer. The argmax over c returns (a+b) mod P by constructive interference of waves, not by table lookup.
A worked micro-example
Shrink P to 5 with a single key frequency w = 2π/5 to see the peak form by hand. Take a = 2, b = 4, so the answer is (2+4) mod 5 = 1. The readout score for each candidate c is cos(w·(6 − c)):
c = 0 : cos(432°) ≈ +0.31
c = 1 : cos(360°) = +1.00 ← max
c = 2 : cos(288°) ≈ +0.31
c = 3 : cos(216°) ≈ −0.81
c = 4 : cos(144°) ≈ −0.81The score peaks unambiguously at c = 1, the correct residue, and it did so without ever computing ‘6’ as an integer or storing the pair (2, 4). A single frequency leaves runners-up (c = 0 and c = 2 tie here); the real model stacks five or six frequencies so their agreement collapses onto one class and the ties vanish. That redundancy is why the circuit is robust and generalizes.
How they knew: ablations and direct evidence
Claiming a network runs an algorithm is cheap; proving it is the hard part, and the modular-addition work is convincing because it triangulates from several angles. The Fourier spectrum of the embeddings is measurably sparse. Individual neurons are visibly periodic in (a, b). Direct logit attribution shows the logits are well approximated by the Σ_k cos(w_k(a+b−c)) formula.
The decisive test is ablation. Project the activations or logits onto the Fourier basis and zero out components. Keep only the key frequencies and the model still works; remove them and it fails. That causal knife — performance depends on these frequencies and not the rest — is what separates a mechanistic account from a suggestive correlation. And those same ablations, run at every training step rather than only at the end, become the progress measures that explain the timing of the grok.