What , '’': depth’ actually means
A transformer is a stack of L identical blocks. Each block takes a sequence of hidden vectors X: [N, d] — N tokens, each a d-dimensional vector — and returns a sequence of the same shape. Depth is L, the count of these blocks; width is d, the size of each vector. Depth scaling means increasing L while keeping d (and the head count, and the FFN expansion ratio) fixed.
Because the input and output shapes of a block are identical, the blocks compose like function iteration: h_L = f_L( … f_2( f_1(x) ) ). A token’s representation is refined once per layer, so depth sets how many sequential refinement steps the model can take — the intuition behind depth as ‘reasoning steps.’ A computation that genuinely needs the output of step three before step four cannot be flattened into a wider-but-shallower network. Width gives a layer more room to work in parallel; depth gives the model more steps in series — the distinction running through everything below.
The parameter math: linear in depth
Almost all of a transformer’s non-embedding parameters live inside the repeated block, and every block is the same size. So counting is easy. One block holds two big pieces. Attention has four d × d projection matrices — W_Q, W_K, W_V, W_O — contributing 4 · d^2 parameters. The FFN has two matrices, d × d_ff and d_ff × d; with the standard d_ff = 4d that is 2 · d · 4d = 8 · d^2. Per block, then:
params_per_block ≈ 4·d^2 (attention) + 8·d^2 (FFN)
= 12 · d^2
params_total ≈ L · 12 · d^2 (non-embedding)The whole point of depth scaling lives in that formula. Total parameters are linear in L and quadratic in d. Double the depth and you double the block parameters; double the width and you quadruple them. Depth is the gentle, proportional knob; width is the aggressive one. That asymmetry — O(L) versus O(d^2) — is the single most important fact about the depth axis, and it drives the budget tradeoff at the end.
A worked example
Take a concrete small model: width d = 768, FFN d_ff = 3072 = 4d, and L = 12 layers — roughly a GPT-2-small shape. Per block:
per_block = 12 · d^2 = 12 · 768^2 = 12 · 589,824 ≈ 7.08M
L = 12 → 12 · 7.08M ≈ 85M non-embedding paramsNow scale depth from 12 to 24, holding width fixed. Parameters go to 24 · 7.08M ≈ 170M — exactly double, because the dependence is linear. Compare scaling width instead: keep L = 12 but take d from 768 to 1536. Per block becomes 12 · 1536^2 ≈ 28.3M, so the model jumps to 12 · 28.3M ≈ 340M — a 4× increase for a 2× width. Same headline ‘double a dimension,’ wildly different cost: doubling depth is +85M, doubling width is +255M. When you want a modest parameter bump, adding a few layers is the surgical move; width increases blow up the budget fast.
What extra depth buys
Empirically, adding layers helps — up to a point. More depth lowers loss and improves tasks that reward multi-step, compositional computation: tracking long-range dependencies, resolving nested structure, chaining intermediate inferences. The mechanism is the serial one from earlier: each layer reads what previous layers wrote into the residual stream and builds on it, so a k-step computation needs on the order of k layers to unfold. A shallow-but-wide model has huge per-step capacity but few steps, and struggles with problems whose answer depends on its own intermediate results.
Picture the residual stream as a shared workspace every layer reads from and writes to; depth is how many times the model revises it before committing to an answer. That predicts both the benefit and the ceiling: once the computation has converged, extra passes add little.
Why very deep models are hard to train
Depth’s cost shows up in optimization, not parameter count. Each block is a residual update x → x + F(x). At initialization F(x) is roughly independent noise, so its variance adds to the residual stream at every layer; after L layers the activation variance has grown by a factor of about L. The signal entering the final LayerNorm can be an order of magnitude larger in a 100-layer net than a 10-layer one, which drives saturated activations and unstable early training.
Gradients face the mirror-image problem on the way back. The chain rule multiplies a Jacobian per layer; across many layers those factors compound, so gradients can decay toward zero (vanishing) or blow up (exploding) before they reach the earliest blocks. The deeper the stack, the longer this multiplicative chain and the more fragile the signal — which is why naive deep post-norm transformers refuse to train past a few dozen layers without help.
Fix 1: pre-norm and a clean residual path
The first and most important fix is where you put the LayerNorm. The original transformer used post-norm: x → LayerNorm(x + F(x)). The normalization sits on the residual path, so the identity shortcut is repeatedly rescaled and the clean gradient highway is broken. Modern deep models use pre-norm: x → x + F(LayerNorm(x)). Here the normalization is applied only inside the sublayer’s branch, and the residual path from input to output is a pure sum of identities.
That pure additive path is what tames deep training. Because the shortcut is an unmodified identity, gradients flow from the loss straight back to layer one without being multiplied down at every step — the vanishing-gradient chain is broken by construction. Pre-norm transformers are far more stable at depth and tolerate larger learning rates with less warmup, which is why essentially every large model since GPT-2 is pre-norm. It can slightly under-use the deepest layers, but the stability it buys is decisive.
Fix 2: residual scaling and DeepNet-style init
Pre-norm alone still lets forward variance grow with depth, so the second family of fixes attacks that directly by shrinking each residual contribution. If every branch is scaled by 1/√(2L) (or the output projections are initialized proportionally smaller, as GPT-2 does with its 1/√N scaling), the summed variance across L layers stays bounded near one instead of growing like L. The model starts close to the identity function and grows its effective depth gradually as training proceeds.
DeepNet makes this principled. Its DeepNorm scheme scales the residual branch by a constant α > 1 before the addition and initializes the sublayer weights with a matching factor β < 1, both chosen as functions of L so the expected update per step is bounded regardless of depth. That bound is what let DeepNet train 1000-layer transformers stably, where ordinary post-norm diverges immediately. The common thread: keep the per-layer perturbation small so a long stack behaves, at initialization, like a shallow one.