← All Posts
Deep Learning · Transformers· Reading paths and labs

The Transformer Evolution Map: Compare Changes on the Right Axis

Begin with a complete baseline

GPT-2 supplies a concrete causal decoder: learned absolute positions, multi-head softmax attention, a dense GELU MLP, LayerNorm, and residual additions. Training predicts the next token. Use this as a reference graph, not as a claim that every later model descends through one mandatory sequence of modifications.

A map of design axes

QuestionDesign choicesWhat changes?
How can tokens exchange information?Softmax attention, sparse/window attention, kernel recurrence, KDAConnectivity or token-mixing function
What attention history is stored?MHA / MQA / GQA, MLAHead sharing or latent representation
How does position enter?Additive positions, RoPE, ALiBi, order-sensitive recurrencePosition-dependent features or scores
Which feed-forward parameters run?Dense/gated MLP, MoENonlinear transformation and conditional capacity
How are depth contributions combined?Residual addition, AttnResInputs received from earlier depth sources
How is the operation executed?FlashAttention, paging, distributed ring, speculationMemory schedule, allocation, distribution, or verified generation schedule

Selected milestones, with separate branches

The 2017 Transformer establishes attention-based encoder-decoder modeling. BERT and GPT-2 illustrate different context masks and objectives. Subsequent work branches into representation, state, and execution changes rather than a single ladder where every new item supersedes the last.

On the execution branch, FlashAttention changes the IO schedule of exact attention. On the state branch, linear Transformers express kernel attention as a recurrence. DeltaNet introduces error-correcting state writes, Gated DeltaNet adds scalar forgetting, and KDA uses more granular decay.

On the cache-representation branch, head sharing and latent compression reduce how much token-indexed information is stored. Hybrid stacks combine recurrent and attention layers. Kimi K3 is a concrete composition of these ideas with latent MoE and block AttnRes, described using its July 2026 report and pinned configuration.

Which comparisons are meaningful?

“FlashAttention versus GQA” is not a clean either/or choice: GQA defines head sharing, and a compatible fused kernel can execute it. “MoE versus linear attention” also mixes axes: one changes the feed-forward function and the other the token mixer. A model can use both.

“Softmax attention versus a delta recurrence” is a more direct mixer comparison, but it still requires matched budgets and tasks. Full attention retains token-indexed history; a fixed recurrent state compresses it. Evaluate exact retrieval, competing associations, sequence length, quality, and measured execution cost instead of declaring a universal winner from asymptotic notation.

A fair comparison sheet

Record training tokens and data policy, total and active parameters, tokenizer, context length, precision, hardware, batch size, and kernel implementation. Report prefill latency, decode latency or time per output token, throughput, and persistent-state memory separately. Use the same prompts and decoding policy when comparing generated outputs.

For an algebraically equivalent implementation change, first verify outputs and, where applicable, gradients within justified tolerance. For an architectural change, retraining or adaptation is part of the comparison: a shape-compatible substitution is not automatically function-preserving.