GPT-2 as a Starting Point: Build the Whole Language Model
Fix one concrete baseline
The GPT-2 report describes a causal decoder language model with learned absolute positions, layer normalization before sublayers, and a final layer normalization. For a reproducible implementation reference, inspect the released model code and the 124M hyperparameters.
We will use vocabulary $V=50{,}257$, context table size 1024, residual width $d=768$, 12 layers, 12 attention heads, and MLP width $4d=3072$. The per-head Q/K/V width is 64 in this baseline. These are GPT-2-specific values, not requirements imposed on every decoder.
Read the forward pass in five moves
- Token IDs select rows from a $V\times d$ embedding table.
- Absolute position IDs select rows from a $1024\times d$ position table; add the two vectors at each position.
- Apply 12 pre-norm residual blocks, each containing causal multi-head attention followed by a GELU MLP.
- Apply the final layer normalization.
- Multiply by the transpose of the token embedding table to obtain $V$ logits per position.
The tied vocabulary projection reuses token embedding parameters; it does not add another independent $V\times d$ matrix. The logits at the last input position predict the next token. During training, earlier positions also contribute their shifted next-token losses.
h = token_embedding[token_ids] + position_embedding[position_ids]
for block in blocks:
h = h + block.attention(block.norm1(h), causal_mask)
h = h + block.mlp(block.norm2(h))
h = final_norm(h)
logits = h @ token_embedding.T
This is a graph sketch, not a checkpoint loader: exact projection layout, activation approximation, initialization, dtype, and cache handling must follow the implementation being reproduced.
Count parameters instead of memorizing a model label
Token embeddings contribute $50{,}257\times768=38{,}597{,}376$ parameters and positions contribute $1024\times768=786{,}432$. In one block, Q/K/V and output projection matrices contribute $4d^2$, while the two MLP matrices contribute $8d^2$. That is $12d^2=7{,}077{,}888$ matrix parameters per block.
Include four attention projection biases totaling $4d$, MLP biases totaling $5d$, and two layer norms with scale and offset totaling $4d$. The block total is therefore $12d^2+13d=7{,}087{,}872$. Twelve blocks contribute 85,054,464. The final layer norm adds $2d=1536$.
This is a parameter-accounting derivation for the stated architecture. Model-size labels are rounded, and different repositories sometimes count or name models differently. The explicit inventory makes discrepancies diagnosable.
Why plain text supplies targets
For a document, every observed next token becomes a supervised target for its preceding prefix. The objective is the sum of negative log conditional probabilities. No separate human label is required for each next-token example, although data selection and formatting strongly influence what the model learns.
Suppose a training fragment is Question: ... Answer: .... The model can learn the conditional distribution of answer text because those tokens appear after the question in its sequence. Whether it reliably follows novel instructions is an empirical capability question, not a special architectural operation named “instruction following.”
A prompt is another prefix
Encode the prompt, run prefill, and use the final logits with a chosen decoding policy. Feed the sampled token back at the next absolute position. Stop according to EOS or an explicit length limit. GPT-2's learned position table has a finite indexed range; increasing a configuration number alone does not train meaningful new position embeddings.
For efficient decoding, cache each layer's keys and values. This baseline uses one K/V head for each query head, so cache bytes grow with sequence length, layer count, and width. The KV-cache guide develops the exact state and reuse conditions.
Change one component at a time
Replace learned absolute positions with RoPE to change how order enters attention. Replace LayerNorm with RMSNorm or GELU MLPs with gated MLPs to change the block. Use GQA or MLA to change cached representation. Replace selected attention layers with KDA to introduce recurrent state. Use MoE to activate different feed-forward parameters per token. Use AttnRes to change depth aggregation.
These substitutions are not all algebraically equivalent optimizations. Most change the learned model and require training or adaptation. FlashAttention, in contrast, implements the same attention operation with a different memory schedule, subject to floating-point effects. That distinction is the organizing principle of the evolution map.
Try it: Why does tying the output projection not eliminate the vocabulary matrix multiplication?
Tying saves a separate parameter matrix. Computing logits still multiplies hidden states by the transposed embedding table, so the vocabulary projection still requires arithmetic.
Move from a baseline to design tradeoffs
Continue with the quadratic problem and execution phases, then follow the comparison map. The site also has a separate from-scratch Transformer implementation track.