← All Posts
Deep Learning · Diffusion & Flow Models · MIT 6.S184· Lecture 3

Flow Matching, Score Matching & Diffusion Models

The goal: turn Lecture 2's formulas into a model we can train. We can make a noisy example and calculate a useful label for it. Train a network to predict that label, then use its predictions to move fresh noise into a generated sample.

What this lecture adds

Lecture 2 chose a sequence of distributions from noise to data and derived the velocity that follows it. Here we learn that velocity from examples. We also learn an alternative, the score, and show how it gives the velocity needed for generation.

The route is: make training examples → learn velocity or score → generate with an ODE or SDE. Each step below explains what we know, what the network predicts, and why the formula works.

The training example and its symbols

Time runs from noise at $t=0$ to data at $t=1$. Pick a data example $Z$ and independent Gaussian noise $\varepsilon$, then mix them:

$$Z\sim p_{\mathrm{data}},\qquad \varepsilon\sim\mathcal N(0,I_d),\qquad X_t=\alpha_tZ+\beta_t\varepsilon.$$

$\alpha_t$ controls how much data we include; $\beta_t$ controls the noise scale. They start at $(0,1)$ and end at $(1,0)$, giving $X_0=\varepsilon$ and $X_1=Z$. Here $d$ is the number of coordinates; $I_d$ means the noise has independent coordinates of variance one. Capital letters denote random samples; lowercase $x,z$ denote particular values.

NotationMeaning here
$p_t(x\mid z)$Density of noisy examples made from one fixed data example $z$; a Gaussian with mean $\alpha_tz$ and covariance $\beta_t^2I_d$.
$p_t(x)$Density after mixing over all data examples and hiding which $z$ was used. This is the marginal density.
$u_t(x)$; $u_t^\theta(x)$The desired velocity; the network's predicted velocity. $\theta$ denotes its trainable parameters.
$\dot\alpha_t$, $\nabla_x$$\dot\alpha_t=d\alpha_t/dt$. The gradient $\nabla_x$ differentiates with respect to each coordinate of $x$, holding time fixed.

Throughout the losses, $\mathbb E$ means average over the sampled training examples, and $\|\text{prediction}-\text{label}\|^2$ means sum the squared errors across coordinates. Training estimates this average with a batch. Unless stated otherwise, sample $t$ uniformly; score formulas need $\beta_t>0$ and the time weighting discussed below.

1. Flow matching: predict a velocity

What should the model learn? Given only $(t,x)$, output the velocity $u_t(x)$ that moves the full distribution along our chosen path. Ideally we would minimize:

$$\mathcal L_{\mathrm{FM}}(\theta)=\mathbb E_{t,X_t}\left[\|u_t^\theta(X_t)-u_t(X_t)\|^2\right].$$

The problem is the label $u_t(X_t)$: evaluating it requires averaging over all possible hidden data examples. We can instead compute the velocity for the particular pair $(Z,\varepsilon)$ we sampled. Hold that pair fixed and differentiate its position:

$$\underbrace{u_t(X_t\mid Z)}_{\text{conditional velocity label}}=\frac{dX_t}{dt}=\frac d{dt}(\alpha_tZ+\beta_t\varepsilon)=\dot\alpha_tZ+\dot\beta_t\varepsilon.$$

This gives conditional flow matching (CFM): use the same network, with a label we can calculate.

$$\boxed{\mathcal L_{\mathrm{CFM}}(\theta)=\mathbb E_{t,Z,\varepsilon}\left[\|u_t^\theta(\alpha_tZ+\beta_t\varepsilon)-(\dot\alpha_tZ+\dot\beta_t\varepsilon)\|^2\right].}$$

Input: time and the noisy point. Label: the conditional velocity. We use $Z$ and $\varepsilon$ to build the example, but do not give them to the network. Here “conditional” refers to the known training endpoint, not a text prompt.

The simplest choice: a straight line

Choose $\alpha_t=t$ and $\beta_t=1-t$, so $\dot\alpha_t=1$ and $\dot\beta_t=-1$:

$$X_t=(1-t)\varepsilon+tZ,\qquad \frac{dX_t}{dt}=Z-\varepsilon.$$

The input changes with time, but the velocity label stays constant for a fixed pair. One training step is:

  1. Sample a batch of independent $(t,Z,\varepsilon)$ and construct each $X_t$.
  2. Predict $u_t^\theta(X_t)$ and compare it with the label $Z-\varepsilon$.
  3. Average $\|u_t^\theta(X_t)-(Z-\varepsilon)\|^2$ over the batch, then update $\theta$ by gradient descent.

Training does not require solving an ODE.

One training pair. Move time to see the input move along its line. Changing noise or data chooses a different pair; these sliders do not simulate a trained model.

Schedules
$\alpha_t=t,\quad\beta_t=1-t$

Input position
$X_t=(1-t)\varepsilon+tz$

Velocity label
$u_t(X_t\mid z)=dX_t/dt=z-\varepsilon$

Blue: starting noise. Orange: data endpoint. Green: the current training input. The line's slope is the velocity label; moving time along a fixed line leaves its slope unchanged.

Each training pair follows a straight line. Generated paths can still curve: the learned velocity averages suggestions from many possible pairs, and those averages change as the sample moves. The lecture calls this a conditional optimal-transport (CondOT) path; that name does not mean the independently sampled noise–data pairs give a globally optimal matching between distributions.

2. Why those labels teach the right velocity

Different hidden pairs can produce the same observed $(t,x)$ and suggest different velocities. Squared-error training learns their average, weighted by how likely each pair is after observing $(t,x)$. That is exactly the marginal velocity from Lecture 2:

$$u_t(x)=\mathbb E\left[u_t(X_t\mid Z)\mid t,X_t=x\right].$$

The part after the bar means “given this time and this observed point.” The expectation averages the remaining uncertainty about $Z$. It uses the posterior weights from Lecture 2, so equally frequent data examples need not get equal weight at a particular $x$.

More precisely, the two losses differ only by a nonnegative constant that the network cannot change:

$$\boxed{\mathcal L_{\mathrm{CFM}}(\theta)=\mathcal L_{\mathrm{FM}}(\theta)+C_{\mathrm{flow}}.}$$

Consequently they have the same expected parameter gradients and the same best-fitting network. A perfect marginal prediction can still have positive conditional loss, because the individual labels disagree.

Proof: expand the squared error and see why the cross term is zero

Write $U=u_t(X_t\mid Z)$ for the label and $m(t,X_t)=\mathbb E[U\mid t,X_t]$ for its average. The leftover $\eta=U-m(t,X_t)$ has average zero when $(t,X_t)$ is fixed. Let $a=u_t^\theta(X_t)-m(t,X_t)$ be the network's error relative to that average.

$$\|u_t^\theta(X_t)-U\|^2=\|a-\eta\|^2=\|a\|^2-2a^\top\eta+\|\eta\|^2.$$

$a^\top\eta$ is the dot product. For a fixed network and fixed $(t,X_t)$, $a$ is fixed, so averaging $a^\top\eta$ gives $a^\top\mathbb E[\eta\mid t,X_t]=0$. Average the remaining terms:

$$\mathcal L_{\mathrm{CFM}}=\mathcal L_{\mathrm{FM}}+\underbrace{\mathbb E\|\eta\|^2}_{C_{\mathrm{flow}}\ge0}.$$

The constant depends on the training distribution, not $\theta$, so its parameter derivative is zero. This assumes finite squared errors and permission to differentiate under the expectation. Individual batch gradients can differ. A limited network may not represent the exact field, but both expected losses still select the same best fit within that network class.

3. Score matching: predict how density changes with position

Why learn a score? It supplies the correction needed when generation adds random motion. For our Gaussian path, it also lets us recover the velocity, so we can train one network and use either sampler.

$$s_t(x)=\nabla_x\log p_t(x).$$

This differentiates log density with respect to position, holding time fixed. In one dimension, a positive score means log density increases to the right; a negative score means it increases to the left. It is a direction of increasing density, not itself a velocity. A zero score can occur at a peak or a valley.

As before, the ideal label is hard to compute for the whole data mixture. The ideal score-matching (SM) loss is:

$$\mathcal L_{\mathrm{SM}}(\theta)=\mathbb E_{t,X_t}\left[\|s_t^\theta(X_t)-s_t(X_t)\|^2\right].$$

Derive a label we can calculate

For one fixed endpoint, $p_t(x\mid z)=\mathcal N(x;\alpha_tz,\beta_t^2I_d)$. Its density contains $\exp(-\|x-\alpha_tz\|^2/(2\beta_t^2))$. Taking the log removes the exponential; the normalization becomes a term $C$ independent of $x$:

$$\log p_t(x\mid z)=C-\frac{\|x-\alpha_tz\|^2}{2\beta_t^2}.$$

Differentiate with respect to $x$. The derivative of $C$ is zero. The squared term gives $2(x-\alpha_tz)$, just as the derivative of $(x-c)^2$ is $2(x-c)$:

$$s_t(x\mid z)=\nabla_x\log p_t(x\mid z)=-\frac{2(x-\alpha_tz)}{2\beta_t^2}=-\frac{x-\alpha_tz}{\beta_t^2}.$$

For our sampled point, $X_t=\alpha_tZ+\beta_t\varepsilon$, subtracting the mean leaves $\beta_t\varepsilon$. Therefore:

$$s_t(X_t\mid Z)=-\frac{\beta_t\varepsilon}{\beta_t^2}=-\frac{\varepsilon}{\beta_t},\qquad\beta_t>0.$$

The conditional score points back toward its Gaussian mean. Its one-dimensional slope is $-1/\beta_t^2$: smaller noise means a narrower cloud and a stronger score for the same displacement from the mean. At zero noise the conditional law is a point mass, so these density derivatives no longer apply.

4. Denoising score matching: use the known noise as a label

We know the noise used to construct each training example, so the score label above is available. Replace the unavailable marginal score in the loss with this conditional label:

$$\boxed{\mathcal L_{\mathrm{DSM}}(\theta)=\mathbb E_{t,Z,\varepsilon}\left[\left\|s_t^\theta(X_t)+\frac{\varepsilon}{\beta_t}\right\|^2\right].}$$

The plus sign comes from subtracting the negative label: prediction minus $(-\varepsilon/\beta_t)$. This is denoising score matching (DSM): learn from noisy examples how to point toward higher density.

Why averaging the training labels gives the mixture's score

The network sees only $(t,x)$. Different hidden pairs $(Z,\varepsilon)$ can produce that same input, with different labels $-\varepsilon/\beta_t$. As in flow matching, squared-error training learns their average, weighting each possible pair by how likely it is given this input. We must show that this average is the score we want.

Start with the mixture density: average the conditional densities over data examples.

$$p_t(x)=\int p_t(x\mid z)p_{\mathrm{data}}(z)\,dz.$$

Its score is the derivative of its log. Use $\nabla_x\log p_t=(\nabla_xp_t)/p_t$, then differentiate the mixture inside the integral:

$$s_t(x)=\frac{\nabla_xp_t(x)}{p_t(x)}=\int\frac{\nabla_xp_t(x\mid z)}{p_t(x)}p_{\mathrm{data}}(z)\,dz.$$

To expose the conditional score, multiply and divide the integrand by $p_t(x\mid z)$. Nothing changes in value; the two resulting fractions have useful meanings:

$$\begin{aligned} s_t(x)&=\int\underbrace{\frac{\nabla_xp_t(x\mid z)}{p_t(x\mid z)}}_{\text{conditional score }s_t(x\mid z)}\;\underbrace{\frac{p_t(x\mid z)p_{\mathrm{data}}(z)}{p_t(x)}}_{\text{posterior weight }p_t(z\mid x)}\,dz\\ &=\int s_t(x\mid z)\,p_t(z\mid x)\,dz. \end{aligned}$$

The second fraction is Bayes' rule: it weights each endpoint $z$ by how likely it is to have produced the observed $x$. Thus the mixture's score is a weighted average of conditional scores. An endpoint whose cloud is denser at $x$ gets more weight when the prior weights are equal.

Finally, we already derived $s_t(X_t\mid Z)=-\varepsilon/\beta_t$. Replace each conditional score in that average by its noise-based label:

$$\begin{aligned} s_t(x)&=\mathbb E[s_t(X_t\mid Z)\mid t,X_t=x]\\ &=\mathbb E[-\varepsilon/\beta_t\mid t,X_t=x]. \end{aligned}$$

The bar means keep time and the observed point fixed; average over the hidden data–noise pairs that could have produced that point. This is exactly the average learned by squared-error training, so the learned target is the marginal score.

The same loss proof therefore gives $\mathcal L_{\mathrm{DSM}}=\mathcal L_{\mathrm{SM}}+C_{\mathrm{score}}$, when the losses are finite. Here $C_{\mathrm{score}}$ is the average squared difference between an individual conditional label and the marginal score:

$$C_{\mathrm{score}}=\mathbb E_{t,Z,\varepsilon}\left[\left\|-\frac{\varepsilon}{\beta_t}-s_t(X_t)\right\|^2\right]\ge0.$$

Different hidden pairs can give different labels for the same $(t,x)$; this constant measures their disagreement around their average. It contains no network prediction or parameters $\theta$, so training cannot change it. Even a network that predicts the exact marginal score has DSM loss $C_{\mathrm{score}}$. The density calculation above assumes we can differentiate inside the integral.

Predict noise to avoid a label that grows as noise shrinks

Dividing by small $\beta_t$ makes the score label large. Instead, let a network $\varepsilon_t^\theta(x)$ predict the noise, then convert its prediction to a score with $s_t^\theta(x)=-\varepsilon_t^\theta(x)/\beta_t$. Substitution into the squared error gives:

$$\left\|s_t^\theta(X_t)+\frac{\varepsilon}{\beta_t}\right\|^2=\frac{1}{\beta_t^2}\|\varepsilon_t^\theta(X_t)-\varepsilon\|^2.$$

Multiplying each example's score loss by $\beta_t^2$ cancels that factor. We get ordinary noise-prediction mean squared error:

$$\mathcal L_{\mathrm{noise}}(\theta)=\mathbb E_{t,Z,\varepsilon}\left[\|\varepsilon_t^\theta(X_t)-\varepsilon\|^2\right].$$

The training procedure is the flow-matching procedure with a different label: predict $\varepsilon$ instead of $Z-\varepsilon$. The best noise prediction is $\mathbb E[\varepsilon\mid t,X_t]$. This is why “predict the noise” does not mean the model can identify the exact noise in every example.

Time weighting and the zero-noise endpoint

Positive weights depending only on time preserve the ideal target at each interior time, but change which times dominate training; a finite network's best fit can therefore change. Ordinary noise MSE corresponds to weighted DSM, not unweighted DSM.

For $\beta_t=1-t$, the score label has expected squared size $\mathbb E\|\varepsilon/\beta_t\|^2=d/(1-t)^2$. This grows so quickly near $t=1$ that its uniform time average diverges. Merely avoiding the exact endpoint does not ensure finite score losses; use a cutoff or suitable weighting before applying the finite-loss proof. Noise prediction keeps the label's expected squared size at $d$, although converting back to a score still divides by $\beta_t$.

5. Turn a score prediction into a velocity

Purpose: sampling needs a velocity, but we may have trained a score or noise network. Both labels come from the same Gaussian construction, so we can express one using the other.

Work at a time with $\alpha_t\ne0$ and $\beta_t>0$, since the derivation divides by these quantities. For one training pair, abbreviate its conditional score as $s=s_t(x\mid z)$. First solve for the hidden noise, then for the endpoint:

$$s=-\frac{\varepsilon}{\beta_t}\ \Longrightarrow\ \varepsilon=-\beta_ts,$$
$$x=\alpha_tz+\beta_t\varepsilon=\alpha_tz-\beta_t^2s\ \Longrightarrow\ z=\frac{x+\beta_t^2s}{\alpha_t}.$$

Substitute these into the conditional velocity $\dot\alpha_tz+\dot\beta_t\varepsilon$:

$$\begin{aligned} u_t(x\mid z)&=\dot\alpha_t\frac{x+\beta_t^2s}{\alpha_t}-\dot\beta_t\beta_ts\\ &=\frac{\dot\alpha_t}{\alpha_t}x+\left(\frac{\dot\alpha_t}{\alpha_t}\beta_t^2-\beta_t\dot\beta_t\right)s. \end{aligned}$$

The equation above still assumes we know the endpoint $z$. During generation we only have the current time and position, so we average over the endpoints that could explain them. Use the posterior weights $p_t(z\mid x)$: how likely each endpoint is after observing $x$ at time $t$. These weights integrate to one.

In this average, only $z$ varies. We hold $t$ and $x$ fixed, so $\alpha_t$, $\beta_t$, and their time derivatives are fixed too. The two multipliers in the conversion therefore come outside the integral:

$$\begin{aligned} u_t(x)&=\int u_t(x\mid z)\,p_t(z\mid x)\,dz\\ &=\frac{\dot\alpha_t}{\alpha_t}x\underbrace{\int p_t(z\mid x)\,dz}_{1}\\ &\quad+\left(\frac{\dot\alpha_t}{\alpha_t}\beta_t^2-\beta_t\dot\beta_t\right)\underbrace{\int s_t(x\mid z)\,p_t(z\mid x)\,dz}_{s_t(x)}. \end{aligned}$$

The first line is the marginal velocity's averaging rule from Section 2. In the second line, averaging the same value $(\dot\alpha_t/\alpha_t)x$ over weights that sum to one leaves that value unchanged. In the last line, the weighted average of conditional scores is exactly the marginal score, as proved in Section 4. This gives:

$$\boxed{u_t(x)=\frac{\dot\alpha_t}{\alpha_t}x+\left(\frac{\dot\alpha_t}{\alpha_t}\beta_t^2-\beta_t\dot\beta_t\right)s_t(x).}$$

For the straight-line schedules

Substitute $\alpha_t=t$, $\dot\alpha_t=1$, $\beta_t=1-t$, and $\dot\beta_t=-1$. The score coefficient simplifies to $(1-t)^2/t+(1-t)=(1-t)/t$:

$$u_t(x)=\frac{x}{t}+\frac{1-t}{t}s_t(x),\qquad 0<t<1.$$

If the network predicts noise, substitute $s_t^\theta=-\varepsilon_t^\theta/(1-t)$; the factors $1-t$ cancel:

$$u_t^\theta(x)=\frac{x-\varepsilon_t^\theta(x)}{t}.$$

Going the other way, multiply the score–velocity equation by $t$, subtract $x$, then divide by $1-t$:

$$s_t^\theta(x)=\frac{tu_t^\theta(x)-x}{1-t}.$$

These are conversions between network outputs; they do not require knowing the endpoint at generation time. They also involve divisions: score-to-velocity needs $t>0$, while velocity-to-score needs $t<1$.

General inverse and why convertible outputs can have different losses

To shorten the general identity, call its coefficients $A_t=\dot\alpha_t/\alpha_t$ and $B_t=A_t\beta_t^2-\beta_t\dot\beta_t$. Then $u_t=A_tx+B_ts_t$, so $s_t=(u_t-A_tx)/B_t$ when $\alpha_t\ne0$ and $B_t\ne0$.

If predictions are tied by this identity, subtracting the corresponding conditional labels cancels $A_tx$:

$$\|u_t^\theta(X_t)-u_t(X_t\mid Z)\|^2=B_t^2\|s_t^\theta(X_t)-s_t(X_t\mid Z)\|^2.$$

Flow matching is thus a time-weighted score loss. On the straight-line path, $B_t=(1-t)/t$, giving noise-prediction error weighted by $1/t^2$. Algebraically related outputs do not make their unweighted training losses identical.

6. Use the learned quantities in a diffusion model

In this lecture, a denoising diffusion model uses a Gaussian probability path. The conversion above means its velocity and score can come from one trained network.

The ODE moves samples with velocity $u_t$. To allow random motion while keeping the same distributions at each time, Lecture 2 adds both fresh noise and a score correction:

$$dX_t=\underbrace{\left[u_t(X_t)+\frac{\sigma_t^2}{2}s_t(X_t)\right]dt}_{\text{directed movement}}+\underbrace{\sigma_t\,dW_t}_{\text{random movement}}.$$

$W_t$ is Brownian motion. The chosen coefficient $\sigma_t$ controls noise added during sampling; $\beta_t$ describes the noise already present in the chosen probability path. They have different jobs.

Why add the score term? Random movement spreads probability out. The score adds movement toward increasing density, counteracting that extra spreading. With the exact fields and correct starting distribution, the resulting SDE follows the same densities $p_t$ as the ODE. Setting $\sigma_t=0$ recovers the ODE.

Generate a new sample, one step at a time

Start with fresh $x_0\sim\mathcal N(0,I_d)$. For a velocity network, split time into $N$ steps of width $h=1/N$, with $t_k=kh$. Euler sampling uses “new position = old position + velocity × elapsed time”:

$$x_{k+1}=x_k+h\,u_{t_k}^\theta(x_k),\qquad k=0,\ldots,N-1.$$

Return $x_N$ as the generated example. We choose no data endpoint $Z$ during this process; the learned field guides the sample. More, smaller steps reduce solver error under suitable smoothness, but cannot fix an inaccurate network.

For stochastic sampling, add the score correction and draw fresh independent $\xi_k\sim\mathcal N(0,I_d)$ each step:

$$x_{k+1}=x_k+h\left[u_{t_k}^\theta(x_k)+\frac{\sigma_{t_k}^2}{2}s_{t_k}^\theta(x_k)\right]+\sigma_{t_k}\sqrt h\,\xi_k.$$

The $\sqrt h$ comes from Brownian increments: their variance over a time interval $h$ is $h$, so their standard deviation is $\sqrt h$. This update is called Euler–Maruyama.

Where endpoint divisions need care

A directly trained velocity network can start at $t=0$ without a conversion. A noise-to-velocity conversion divides by $t$ and cannot be evaluated there. Starting it at $t=\delta>0$ avoids division by zero, but exact sampling would then need the distribution $p_\delta$; initializing with standard noise is an approximation.

The straight-line velocity-to-score conversion divides by $1-t$. The updates above use $t_k<1$, avoiding the exact terminal division, although errors can still be amplified near the endpoint. Exact distribution preservation describes true fields and a well-defined continuous-time process; learned fields and finite steps introduce approximations.

Keep these training labels together

All three networks receive the same $(t,X_t)$, with $X_t=\alpha_tZ+\beta_t\varepsilon$. The table separates the label for one pair from what squared-error training learns after averaging possible pairs.

PredictLabel for one pairIdeal learned output
Velocity $u_t^\theta$$\dot\alpha_tZ+\dot\beta_t\varepsilon$
Straight line: $Z-\varepsilon$
Marginal velocity $u_t(x)$
Score $s_t^\theta$$-\varepsilon/\beta_t$Marginal score $s_t(x)$
Noise $\varepsilon_t^\theta$$\varepsilon$$\mathbb E[\varepsilon\mid t,X_t=x]$

Where this sits in the literature

The lecture uses MovieGen and Stable Diffusion 3 as flow-matching applications. The training target is one part of a generator; architecture, representation, and inputs such as text prompts are additional choices. See the lecture's application examples.

Check the time direction when comparing formulas: these notes go from noise to data; many diffusion papers start their clock at data. Reversing an ODE changes its velocity's sign. Reversing an SDE also requires a score term, so flipping the drift's sign alone is insufficient.

Beyond Gaussian noise

Flow matching needs a conditional path we can sample and a velocity label we can calculate. Its source need not be Gaussian: the lecture discusses bridges such as low-resolution to high-resolution images, silent video to video with audio, and unperturbed to perturbed cell states. See the bridge examples.

A bridge must specify how source and target examples are paired, called their coupling. Producing plausible high-resolution images alone does not ensure they match a particular low-resolution input; pairing or conditioning must encode that relationship. Different modalities also need a compatible representation. The Gaussian score–velocity conversion above applies to our specific construction, not automatically to every bridge.

Optional checks with short answers

Why can a perfect model still have positive conditional loss?

Different hidden pairs can supply different labels for the same $(t,x)$. The perfect model predicts their average; the remaining squared disagreement is the constant $C$ in the loss proof.

Does the label $Z-\varepsilon$ work for any schedules?

No. Always differentiate $X_t=\alpha_tZ+\beta_t\varepsilon$ to get $\dot\alpha_tZ+\dot\beta_t\varepsilon$. Only the straight-line schedules here have derivatives $1$ and $-1$.

Why is adding random motion alone insufficient?

It spreads probability beyond the chosen path. The score correction cancels that added spreading; without it the distributions generally change.

Sources and reading route

Based on my six-page handwritten notes, Lecture 3 floor matching and diffusion models (14 December 2025): flow training and its proof on pages 1–3; score matching and DSM on pages 3–4; conversion and bridges on pages 5–6. The intermediate algebra, loss weighting, and sampling explanations expand those notes.

Companion material: MIT 6.S184 Lecture 3 slides and the 2025 course page. For the construction of the targets themselves, return to Lecture 2.

Check yourself

Practice: explain the method without looking at the formulas
  • State the network's input and label for flow, score, and noise prediction.
  • Explain why knowing a conditional label is enough to learn a marginal field.
  • Recover velocity from a score prediction using the two substitutions in Section 5.
  • Explain why training can sample any time directly, while generation follows successive steps.

Continue into language: the Large Language Diffusion Models series builds on these probability paths and conditional targets to explain token corruption, masked training, discrete flow matching, and practical text generation.