Transformer Assembly & Weight Initialization
The residual stream
The most useful way to read a modern transformer is not as a stack of layers but as a single residual stream that every block reads from and writes to:
Two properties matter. The stream is never transformed in place, only added to, so there is an unbroken identity path from the embedding to the final norm. And each block sees a normalized copy of the stream while writing an unnormalized increment back into it.
Pre-norm versus post-norm
Post-norm (original)
The normalization sits on the residual path, so every backward pass is multiplied by the norm's Jacobian. Over dozens of layers those factors compound and the gradient reaching early layers is unreliable. Training deep post-norm stacks requires a learning-rate warmup and is still fragile.
Pre-norm (universal today)
The residual path is a clean identity. Differentiating gives $\partial\mathbf x_{\text{out}}/\partial\mathbf x=I+\partial F/\partial\mathbf x$, so the gradient always has an unattenuated route to every layer. This is what makes 80-layer stacks trainable without heroics.
Pre-norm has one wart: the stream itself is never normalized, so its magnitude grows with depth. That is why a final Norm is applied after the last block, before the output projection.
RMSNorm
LayerNorm centres and rescales:
RMSNorm drops the mean subtraction and the bias, keeping only the scale:
Why it is enough
The empirical finding is that the re-centring contributes little; the benefit comes from constraining the magnitude. Removing it loses no measurable quality.
Why it is better
One reduction pass instead of two, no mean to store for the backward pass, and $d$ fewer parameters per norm. On a memory-bound elementwise kernel that is a real saving.
The gated feed-forward network
The classic FFN is $W_2\,\phi(W_1\mathbf x)$. The gated variant splits the first projection into two and uses one branch to modulate the other:
The gate lets the layer suppress or pass each hidden unit as a function of the input, rather than applying a fixed pointwise nonlinearity. It reliably beats an ungated FFN at equal parameter count — which is why $d_{\text{ff}}$ shrinks from $4d$ to $\tfrac83 d$ when the third matrix is added, keeping the budget fixed.
Why naive initialization fails with depth
Here is the failure mode, made quantitative. Model the residual stream as a random vector and each block's output as an independent contribution. If the stream enters a block with variance $v$ and the branch contributes variance $\sigma^2$, then because the two are roughly uncorrelated,
A block contains two residual additions, so an $L$-layer model performs $2L$ of them. Starting from $v_0=1$:
The fix is to make each branch's contribution shrink with depth. Scale the weights of the layers that write into the stream — the attention output projection $W_O$ and the FFN down projection $W_{\text{down}}$ — by
Then $\sigma^2\mapsto\sigma^2/2L$ and the total becomes
which is independent of $L$. This is the GPT-2 initialization rule, and it is why the same recipe transfers unchanged from a 12-layer model to an 80-layer one.
The full recipe
| Tensor | Initialization | Reason |
|---|---|---|
| $W_Q,W_K,W_V$ | $\mathcal N(0,\sigma^2)$ with $\sigma=0.02$, or $\sigma=1/\sqrt d$ | Preserve activation scale through the projection. |
| $W_O$, $W_{\text{down}}$ | same, then multiplied by $1/\sqrt{2L}$ | These write into the residual stream; the depth scaling keeps its variance constant. |
| $W_{\text{gate}},W_{\text{up}}$ | $\mathcal N(0,\sigma^2)$ | They feed the nonlinearity, not the stream. |
| RMSNorm gain $\mathbf g$ | all ones | Start as an identity so the block begins as a pure pass-through. |
| Embedding table | $\mathcal N(0,\sigma^2)$ | Sets $v_0\approx1$ entering the stack. |
| Biases, if present | zero | No reason to break symmetry twice. |
Tied embeddings
The input embedding maps token ids to vectors; the output head maps vectors back to logits. Reusing one matrix for both is weight tying:
It saves $Vd$ parameters, which for a small model with a 128k vocabulary is a large fraction of the total. It also acts as a regularizer, since a token's input and output representations are forced to agree. Large models increasingly untie the two, because $Vd$ stops being a meaningful share of the budget and the extra freedom helps slightly.
The assembled block
After $L$ of these, a final RMSNorm and the output projection produce logits. The $1/\sqrt{d_h}$ inside the softmax is the same variance argument in miniature: a dot product of two $d_h$-dimensional unit-variance vectors has variance $d_h$, and dividing by $\sqrt{d_h}$ brings the logits back to order one so the softmax does not saturate.
Pre-norm gives the gradient a clean path, RMSNorm makes normalization cheap, gating makes the FFN more expressive per parameter, and scaling the two stream-writing projections by $1/\sqrt{2L}$ makes the whole recipe depth-independent. That last item is the one that silently breaks when you scale up.
Check yourself
- Derive $v_{2L}=1+2L\sigma^2$ and state the independence assumption it relies on. derivation
- Show that scaling $W_O$ and $W_{\text{down}}$ by $1/\sqrt{2L}$ makes the output variance independent of $L$. derivation
- Explain why pre-norm keeps an unattenuated gradient path but post-norm does not. reasoning
- Count the parameters saved by tying embeddings for $V=128000$, $d=2048$, and express it as a fraction of a 1.5B model. calculation
- Explain the $1/\sqrt{d_h}$ factor using the same variance argument. derivation