The Transformer Evolution Map: Compare Changes on the Right Axis
Begin with a complete baseline
GPT-2 supplies a concrete causal decoder: learned absolute positions, multi-head softmax attention, a dense GELU MLP, LayerNorm, and residual additions. Training predicts the next token. Use this as a reference graph, not as a claim that every later model descends through one mandatory sequence of modifications.
A map of design axes
| Question | Design choices | What changes? |
|---|---|---|
| How can tokens exchange information? | Softmax attention, sparse/window attention, kernel recurrence, KDA | Connectivity or token-mixing function |
| What attention history is stored? | MHA / MQA / GQA, MLA | Head sharing or latent representation |
| How does position enter? | Additive positions, RoPE, ALiBi, order-sensitive recurrence | Position-dependent features or scores |
| Which feed-forward parameters run? | Dense/gated MLP, MoE | Nonlinear transformation and conditional capacity |
| How are depth contributions combined? | Residual addition, AttnRes | Inputs received from earlier depth sources |
| How is the operation executed? | FlashAttention, paging, distributed ring, speculation | Memory schedule, allocation, distribution, or verified generation schedule |
Selected milestones, with separate branches
The 2017 Transformer establishes attention-based encoder-decoder modeling. BERT and GPT-2 illustrate different context masks and objectives. Subsequent work branches into representation, state, and execution changes rather than a single ladder where every new item supersedes the last.
On the execution branch, FlashAttention changes the IO schedule of exact attention. On the state branch, linear Transformers express kernel attention as a recurrence. DeltaNet introduces error-correcting state writes, Gated DeltaNet adds scalar forgetting, and KDA uses more granular decay.
On the cache-representation branch, head sharing and latent compression reduce how much token-indexed information is stored. Hybrid stacks combine recurrent and attention layers. Kimi K3 is a concrete composition of these ideas with latent MoE and block AttnRes, described using its July 2026 report and pinned configuration.
Which comparisons are meaningful?
“FlashAttention versus GQA” is not a clean either/or choice: GQA defines head sharing, and a compatible fused kernel can execute it. “MoE versus linear attention” also mixes axes: one changes the feed-forward function and the other the token mixer. A model can use both.
“Softmax attention versus a delta recurrence” is a more direct mixer comparison, but it still requires matched budgets and tasks. Full attention retains token-indexed history; a fixed recurrent state compresses it. Evaluate exact retrieval, competing associations, sequence length, quality, and measured execution cost instead of declaring a universal winner from asymptotic notation.
A fair comparison sheet
Record training tokens and data policy, total and active parameters, tokenizer, context length, precision, hardware, batch size, and kernel implementation. Report prefill latency, decode latency or time per output token, throughput, and persistent-state memory separately. Use the same prompts and decoding policy when comparing generated outputs.
For an algebraically equivalent implementation change, first verify outputs and, where applicable, gradients within justified tolerance. For an architectural change, retraining or adaptation is part of the comparison: a shape-compatible substitution is not automatically function-preserving.