The FFN you already know: two matrices, one activation
Every transformer layer is an attention sublayer followed by a position-wise feed-forward network (FFN). The original Transformer, and models like BERT and GPT-2/3, use the same simple shape applied independently at every token position:
FFN(x) = activation(x · W1) · W2
x : [d] token vector (model dim d)
W1 : [d, d_ff] up-projection
W2 : [d_ff, d] down-projection
d_ff = 4d the usual expansion ratioThe vector is projected up into a wider hidden space (typically four times the model dimension), a nonlinearity is applied element-wise, and it is projected back down. The activation was ReLU originally and GELU in most later models. Two weight matrices, one activation, no gate. This block holds roughly two-thirds of a transformer’s parameters, so any change to its design — more expressive, or cheaper — pays off across the whole model. SwiGLU is exactly such a change.
The gating idea and the GLU family
A Gated Linear Unit (GLU), introduced by Dauphin et al. in 2016, replaces a single activated projection with a product of two projections, one of which acts as a multiplicative gate:
GLU(x) = (x · W) ⊙ σ(x · V)Here ⊙ is the element-wise (Hadamard) product and σ is the sigmoid. One branch, xW, carries the content; the other branch, σ(xV), produces a value between 0 and 1 for each hidden unit and gates the content — letting some units through fully, damping others toward zero. Crucially, the gate is data-dependent: it is computed from the same input, so the network learns to open and close channels per token. Swap the sigmoid for a different activation and you get the whole GLU family: ReGLU (ReLU gate), GEGLU (GELU gate), and SwiGLU (Swish gate). The sibling article on the GLU family (tm_glu_ffn) treats these variants side by side; here we focus on the Swish one that Llama and PaLM adopted.
Swish / SiLU: the activation underneath
The ‘Swi’ in SwiGLU is Swish, the activation that gates the content branch. Swish is defined as
Swish_β(x) = x · σ(βx) where σ(z) = 1 / (1 + e^(-z))When β = 1 this is called SiLU (Sigmoid Linear Unit), and that is the version Llama and PaLM use. Unlike ReLU, Swish is smooth and non-monotonic: it dips slightly below zero for small negative inputs before recovering, so it does not hard-clip negatives to exactly zero. For large positive x it behaves like the identity (σ(βx) → 1), and for large negative x it decays smoothly toward zero. That smoothness gives well-behaved gradients everywhere — no dead-ReLU regions where the gradient is flat zero — which helps optimization in deep stacks. In SwiGLU, Swish is applied only to the gate branch xW; the value branch xV stays linear.
The SwiGLU FFN: the full formula
Put the gate and the down-projection together and you get the SwiGLU feed-forward block used in production models:
FFN_SwiGLU(x) = ( Swish(x · W) ⊙ (x · V) ) · W2
W : [d, d_ff] gate projection (Swish-activated)
V : [d, d_ff] value projection (linear)
W2 : [d_ff, d] down-projection
biases: usually dropped in modern LLMsRead left to right: the input x is projected two different ways into the hidden space, xW and xV, both of shape [d_ff]. The first is squashed by Swish to become a per-unit gate; the second is left linear. Their element-wise product — the gated hidden state — is then projected back down to [d] by W2. The structural difference from the plain FFN is exactly the extra projection: where GELU-FFN had one up-matrix W1, SwiGLU has two, W and V. That third matrix is what makes gating possible — and what forces the parameter arithmetic in the next section.
Why three matrices — and the 2/3 rule
A plain FFN with expansion d_ff = 4d has two weight matrices, each of size d × d_ff, for a parameter count of
plain FFN: W1 + W2 = 2 · d · d_ff = 2 · d · (4d) = 8 d^2SwiGLU adds the value projection V, so it has three matrices of size d × d_ff: 3 · d · d_ff. If you kept the same d_ff = 4d, SwiGLU would cost 12 d^2 — 50% more parameters than the baseline — and any quality comparison would be unfair. The fix is the 2/3 rule: shrink the hidden width by a factor of two-thirds so the three matrices cost the same as the old two.
match params: 3 · d · d_ff_new = 2 · d · d_ff_old
d_ff_new = (2/3) · d_ff_old = (2/3)(4d) = (8/3) d ≈ 2.667 dSo the SwiGLU hidden dimension is 8/3 d instead of 4d. The gate buys expressiveness; the 2/3 rule pays for it by narrowing the hidden layer, keeping the total parameter and FLOP budget essentially flat.
A worked example: Llama-7B numbers
Take Llama-7B, with model dimension d = 4096. A classic 4× FFN would use d_ff = 4 · 4096 = 16384, at a per-layer cost of
plain FFN: 2 · 4096 · 16384 = 134,217,728 ≈ 134M paramsApplying the 2/3 rule: d_ff = (8/3) · 4096 = 10922.7, which Llama rounds up to a hardware-friendly multiple of 256, giving the actual value d_ff = 11008. The three-matrix SwiGLU cost is then
SwiGLU: 3 · 4096 · 11008 = 135,266,304 ≈ 135M paramsWithin a fraction of a percent of the 134M baseline — the rounding to 256 is the only reason it is not exact. So Llama gets the gating mechanism for effectively free: same parameter count, same compute, but a measurably better loss curve. This is why the seemingly odd hidden size 11008 appears in the Llama config; it is 8/3 × 4096 rounded, not an arbitrary choice.
Why gating improves quality
Why should multiplying two projections beat activating one? The honest answer from the literature (Shazeer’s 2020 GLU Variants Improve Transformer) is largely empirical: across language-modeling benchmarks, GEGLU and SwiGLU consistently edge out ReLU and GELU FFNs at matched parameters. But there is intuition. A plain FFN applies the same nonlinearity to every unit; the gate lets the network apply an input-dependent, per-unit multiplier on top of the content. That multiplicative interaction is more expressive than a fixed point-wise function — it can represent if this feature is present, amplify that one patterns that an additive activation cannot cheaply model. The gate can also suppress irrelevant hidden units to near-zero for a given token, acting like a soft, learned feature selector. Shazeer himself hedged, attributing the gain to ‘divine benevolence,’ but the result replicated widely enough that gated FFNs are now the default in frontier models.