Large Language Diffusion Models: A Reading Roadmap
Start where the previous series ends
The prerequisites are Lecture 2: constructing a training target and Lecture 3: flow matching, score matching, and diffusion, plus familiarity with transformer attention and cross-entropy. We link back to those explanations instead of rebuilding ODEs, SDEs, Gaussian score identities, or transformer blocks.
By the end, you should be able to derive a masked denoising objective, turn a posterior predictor into a sampler, explain why simultaneous token updates can fail, and evaluate a diffusion LLM without confusing iteration count with speed or a likelihood bound with exact perplexity.
Eight chapters, one job each
Carry probability paths and time reversal from MIT 6.S184 into a vocabulary: absorbing masks, transition kernels, and the difference between a denoiser and a joint language model.
Derive the reverse masking posterior, the weighted cross-entropy likelihood bound, and a numerically useful estimator without random masked-count normalization.
Understand why step count changes the generated distribution, distinguish confidence decoding from exact reverse dynamics, and measure parallelization error with an exact toy posterior.
Derive the discrete continuity equation, posterior-averaged jump rates, and the connection between masked diffusion, discrete scores, and flow matching.
Separate the probability model, transformer architecture, training initialization, and inference policy when reading large language diffusion papers.
Turn a denoiser into a conditional generator with correct prompt masking, response supervision, infilling boundaries, and carefully defined guidance.
Reason about repeated full-sequence passes, exact versus approximate KV reuse, block diffusion, likelihood bounds, and fair quality–latency comparisons.
Run an exact reverse-kernel experiment and train a small bidirectional transformer on a synthetic language, with measured results and an analytic denoising-loss reference.
Choose a reading route
For the conceptual bridge: read Parts 1–4 in order. The path becomes a discrete corruption channel, the objective becomes weighted categorical prediction, and the velocity becomes a jump-rate generator.
To start reading model code: read Parts 1–3, then Parts 5–8. Return to Part 4 when a paper parameterizes discrete scores or probability velocities.
To investigate a speed or quality claim: read the finite-step counterexample in Part 3 and the measurement guide in Part 7. Run the lab to see those distinctions in a controlled experiment.
Notation that stays consistent
| Symbol | Meaning |
|---|---|
| $x_0$, $x_t$ | Clean and corrupted sequences; superscript $i$ indexes a token position |
| $m$, $\mathcal V$, $L$ | Input-only mask, clean vocabulary, and sequence length |
| $t$, $a(t)$ | Corruption time and token survival probability; $t=0$ clean, $t=1$ masked |
| $\mu_\theta^i(v\mid x_t,t)$ | Predicted posterior of the clean token at position $i$ |
| $s=1-t$, $\kappa(s)$ | Generative time and reveal schedule, used explicitly in Part 4 |
| $K$ | Number of denoising intervals, distinct from generated token count |
The time direction deliberately differs from the MIT notes until Part 4 switches back. Each derivation states the relevant convention. Unless stated otherwise, examples use fixed-length sequences and a linear, token-independent absorbing mask process.
What the examples establish
The diagrams and interactive examples are explanatory constructions. The toy sampling probabilities are exact calculations. The neural lab results are one recorded CPU experiment. Model-specific statements link to primary papers or official implementations; the model chapter maps foundational design choices rather than tracking the latest leaderboard.
All derivations, diagrams, exercises, and lab code are written for this series. Begin with From Continuous Paths to Token Jumps, or download the lab script if you prefer to learn by running an experiment first.
Research extensions
Continue with these focused notes: