Deep Learning · Transformers· Reading paths and labs
Transformers: From the First Attention Calculation to Hybrid Language Models
Choose your starting point
If Transformers are new, begin with the Foundations model diagram. The same clickable diagram appears in all eight foundation chapters, highlighting the component you are learning and connecting it to the rest of the decoder. If you know Q/K/V but struggle to connect the pieces, follow it through GPT-2 as a complete language model. If you can already implement a decoder, jump to prefill, decode, and complexity.
Chapters define their notation as needed: $n$ is sequence length, $d$ is vector width, and $H$ is the number of attention heads.
A coherent route into modern language models
- Build the reference model. Attention → the residual block → GPT-2 → training.
- Understand growing history. The quadratic problem → KV caching → head sharing → latent attention.
- Replace token history with a recurrent state. Linear attention → DeltaNet → Gated DeltaNet → KDA.
- Compose a modern model. Hybrid attention → MoE → AttnRes → Kimi K3.
- Verify and compare. Use the math lab, then the evolution and comparison map.
All chapters by purpose
Reading paths and labs
- Transformer Foundations: Follow the Model Diagram
- The Transformer Evolution Map: Compare Changes on the Right Axis
- Transformer Math Lab: Test the Identities Behind the Diagrams
Foundations
- Token Embeddings: Just a Table Lookup
- Attention from First Principles: One Complete Calculation
- Multi-Head Attention: Separate Reads, One Residual Update
- Causal Attention: The Information Boundary
- Feed-Forward Networks: Expansion, Nonlinearity, and Gates
- LayerNorm, RMSNorm, and Residual Scale
- A Transformer Block, Line by Line
- Dropout: What Is Random, What Is Preserved
Position and architecture
- Position Encodings: Order Must Enter Somewhere
- RoPE: Relative Position Through Rotated Queries and Keys
- ALiBi: A Distance Bias in Attention Logits
- Cross-Attention: Queries Read a Different Sequence
Attention and memory
- The Quadratic Problem: Prefill, Decode, and a Complexity Explorer
- KV Caching: Why the Prefix Can Stay Put
- MHA, MQA, and GQA: Sharing the KV Cache
- Multi-Head Latent Attention: Derive the Compressed Cache
- Sparse Attention: Windows, Global Tokens, and Information Paths
Recurrent and hybrid models
- Linear Attention: From Pairwise Reads to a Recurrent State
- DeltaNet: Write the Error, Not Another Copy
- Gated DeltaNet: Forget Globally, Update Selectively
- Kimi Delta Attention: Channel-Wise Memory Control
- Hybrid Attention: Combine Retrieval with Recurrent Memory
Architecture and model families
- Transformer Architectures: Choose the Information Flow
- The Transformer Encoder: Context, Padding, and Task Heads
- The Transformer Decoder: From a Prefix to the Next Token
- GPT-2 as a Starting Point: Build the Whole Language Model
- BERT: Learn Bidirectional Representations by Reconstructing Tokens
- Vision Transformers: Turn an Image into a Sequence of Patches
- Mixture of Experts: Conditional MLPs and the Cost of Routing
- Attention Residuals: Let a Layer Choose Its Sources Across Depth
- Kimi K3: Read a Modern Hybrid Model from Its Configuration
Efficient execution
- FlashAttention: Exact Softmax Without the Full Matrix
- FlashAttention-2 and -3: From Algebra to GPU Scheduling
- PagedAttention: Manage the Cache Without Moving the Sequence
- Ring Attention: Distributed Reads with Online Softmax
- Speculative Decoding: Draft, Verify, and Correct