← All Posts
Deep Learning · Transformers· Reading paths and labs

Transformers: From the First Attention Calculation to Hybrid Language Models

Choose your starting point

If Transformers are new, begin with the Foundations model diagram. The same clickable diagram appears in all eight foundation chapters, highlighting the component you are learning and connecting it to the rest of the decoder. If you know Q/K/V but struggle to connect the pieces, follow it through GPT-2 as a complete language model. If you can already implement a decoder, jump to prefill, decode, and complexity.

Chapters define their notation as needed: $n$ is sequence length, $d$ is vector width, and $H$ is the number of attention heads.

A coherent route into modern language models

  1. Build the reference model. Attention → the residual block → GPT-2 → training.
  2. Understand growing history. The quadratic problem → KV caching → head sharing → latent attention.
  3. Replace token history with a recurrent state. Linear attention → DeltaNet → Gated DeltaNet → KDA.
  4. Compose a modern model. Hybrid attention → MoE → AttnRes → Kimi K3.
  5. Verify and compare. Use the math lab, then the evolution and comparison map.

All chapters by purpose

Reading paths and labs

Foundations

Position and architecture

Attention and memory

Recurrent and hybrid models

Architecture and model families

Efficient execution

Training