← All Posts
Deep Learning · Diffusion & Flow Models · Large Language Diffusion Models· Series roadmap

Large Language Diffusion Models: A Reading Roadmap

A continuation of the MIT 6.S184 notes. You already know how continuous probability paths, conditional targets, scores, and numerical sampling fit together. This series transfers that understanding to text: categorical states, masked objectives, parallel decoding, and the design choices behind diffusion language models.

Start where the previous series ends

The prerequisites are Lecture 2: constructing a training target and Lecture 3: flow matching, score matching, and diffusion, plus familiarity with transformer attention and cross-entropy. We link back to those explanations instead of rebuilding ODEs, SDEs, Gaussian score identities, or transformer blocks.

By the end, you should be able to derive a masked denoising objective, turn a posterior predictor into a sampler, explain why simultaneous token updates can fail, and evaluate a diffusion LLM without confusing iteration count with speed or a likelihood bound with exact perplexity.

Eight chapters, one job each

01
From Continuous Paths to Token Jumps

Carry probability paths and time reversal from MIT 6.S184 into a vocabulary: absorbing masks, transition kernels, and the difference between a denoiser and a joint language model.

02
The Masked Diffusion Objective, Derived

Derive the reverse masking posterior, the weighted cross-entropy likelihood bound, and a numerically useful estimator without random masked-count normalization.

03
Sampling: Parallel Reveals, Confidence, and Remasking

Understand why step count changes the generated distribution, distinguish confidence decoding from exact reverse dynamics, and measure parallelization error with an exact toy posterior.

04
Discrete Flow Matching: Probability Moves, Tokens Jump

Derive the discrete continuity equation, posterior-averaged jump rates, and the connection between masked diffusion, discrete scores, and flow matching.

05
LLaDA, Dream, and the Design of a Diffusion LLM

Separate the probability model, transformer architecture, training initialization, and inference policy when reading large language diffusion papers.

06
Conditioning, Infilling, and Guidance

Turn a denoiser into a conditional generator with correct prompt masking, response supervision, infilling boundaries, and carefully defined guidance.

07
Efficiency and Evaluation: What to Measure

Reason about repeated full-sequence passes, exact versus approximate KV reuse, block diffusion, likelihood bounds, and fair quality–latency comparisons.

08
Build a Tiny Diffusion Language Model

Run an exact reverse-kernel experiment and train a small bidirectional transformer on a synthetic language, with measured results and an analytic denoising-loss reference.

Choose a reading route

For the conceptual bridge: read Parts 1–4 in order. The path becomes a discrete corruption channel, the objective becomes weighted categorical prediction, and the velocity becomes a jump-rate generator.

To start reading model code: read Parts 1–3, then Parts 5–8. Return to Part 4 when a paper parameterizes discrete scores or probability velocities.

To investigate a speed or quality claim: read the finite-step counterexample in Part 3 and the measurement guide in Part 7. Run the lab to see those distinctions in a controlled experiment.

Notation that stays consistent

SymbolMeaning
$x_0$, $x_t$Clean and corrupted sequences; superscript $i$ indexes a token position
$m$, $\mathcal V$, $L$Input-only mask, clean vocabulary, and sequence length
$t$, $a(t)$Corruption time and token survival probability; $t=0$ clean, $t=1$ masked
$\mu_\theta^i(v\mid x_t,t)$Predicted posterior of the clean token at position $i$
$s=1-t$, $\kappa(s)$Generative time and reveal schedule, used explicitly in Part 4
$K$Number of denoising intervals, distinct from generated token count

The time direction deliberately differs from the MIT notes until Part 4 switches back. Each derivation states the relevant convention. Unless stated otherwise, examples use fixed-length sequences and a linear, token-independent absorbing mask process.

What the examples establish

The diagrams and interactive examples are explanatory constructions. The toy sampling probabilities are exact calculations. The neural lab results are one recorded CPU experiment. Model-specific statements link to primary papers or official implementations; the model chapter maps foundational design choices rather than tracking the latest leaderboard.

All derivations, diagrams, exercises, and lab code are written for this series. Begin with From Continuous Paths to Token Jumps, or download the lab script if you prefer to learn by running an experiment first.

Research extensions

Continue with these focused notes: