Transformer Foundations: Follow the Model Diagram
Imagine completing “the cat sat …”. Before learning any equations, follow how a transformer turns those three tokens into a prediction. We will open the model a little at a time, keeping the same tokens in view.
Start with the whole model, open its stack of blocks, then look inside one block and its attention heads. The diagram returns in every Foundations chapter, opening at the component you are learning.
Follow token IDs into vectors, watch the blocks update those vectors, and finish with a next-token prediction.
Suppose our tokens are “the cat sat”. A language model uses this prefix to assign a probability to each possible next token.
“on” is an illustrative continuation, not a model measurement. Choosing a token and appending it gives a new prefix: “the cat sat on”.
What happens inside that box? First, turn the tokens into vectors. Then pass those vectors through a stack of blocks.
Embeddings give each token a starting vector. Position information tells the model where it sits. Each block updates the vectors before passing them onward.
Use the last position’s distribution to choose the next token. The final norm and vocabulary head sit after the entire stack.
Now open just one block. Its two jobs are to exchange context across positions and transform the features at each position.
Follow the arrows downward. The side paths carry the input unchanged; each + adds a learned update to it. This running set of vectors is the residual stream.
The “sat” position can read “the”, “cat”, and itself. Each position receives its own contextual update.
The MLP reads the updated U. It applies the same learned function to each position independently.
Y becomes the next block’s X.
This is a pre-norm causal decoder: normalization precedes each branch. Dashed dropout boxes are optional and become identity operations at evaluation. The original 2017 Transformer uses post-norm.
There is one box still to open. Inside self-attention, several heads gather context before their outputs are joined into a single update.
Keep following “sat”. Its query scores the available keys, and the resulting weights combine their value vectors. Every head uses its own learned projections.
The mask applies inside every head. Attention weights are probabilities over context positions; the vocabulary probabilities at the end of the model are over token IDs.
Return to the block. This attention output is added to X at the first +. The MLP then reads the updated U.
Reference: a causal decoder with additive positions and pre-norm blocks. Select a component to read its chapter. Compare encoder and decoder architectures →
From the picture to the numbers
The diagram keeps three token positions all the way through the stack. What changes is the vector at each position: it starts as a lookup from a table, gains context through attention, and is transformed by the MLP. Only the vocabulary head turns these features into scores for possible next tokens.
For now, remember just two sizes: $n$ is the number of tokens and $d$ is the number of features in each token’s vector. The chapters introduce the other symbols when you need them. A block changes the vectors while preserving their $n\times d$ shape.
Eight chapters, one model diagram
The reading order introduces one idea at a time. It differs from execution order: for example, you will study normalization after understanding what its attention and MLP branches do. Use the arrows in the diagram to locate each operation in the forward pass.
- Token embeddings: start at the embedding box. How does an ID select the starting vector for a token?
- One attention head: zoom inside the blue box. How does a query retrieve a weighted combination of values?
- Multi-head attention: zoom back out to the blue box. How do several reads become one update?
- Causal masking: inspect the restriction inside each head. Which token positions may it read?
- The feed-forward network: move to the second branch. How does it transform the features at each position?
- Normalization: find the boxes before both branches and after the stack. How are vectors prepared for the next operation?
- The complete block: follow the continuous path through both + circles. How do the updates combine, and what enters the next block?
- Dropout: inspect the dashed boxes. What changes during training, and what disappears at evaluation?
Then study position information and assemble GPT-2, including its vocabulary head. The diagram uses a causal decoder with normalization before each branch; the architecture chapter explains encoder and encoder–decoder arrangements.
For a complementary visual introduction to the original encoder–decoder model, read Jay Alammar’s The Illustrated Transformer. The diagrams here follow a pre-norm causal decoder, so normalization placement and the decoder’s attention layers differ.
Use the math lab to check the calculations, or the comparison map to explore what changes in later models.