← All Posts
Deep Learning · Diffusion & Flow Models · MIT 6.S184· Lecture 2

Constructing a Training Target: Paths, Fields & Scores

Lecture 2 constructs the target that a generative model should learn. Choose a path of distributions from noise to data, derive a velocity that follows it, then use the score to add stochastic motion without changing those distributions. Lecture 3 explains how conditional regression turns these targets into a trained network.

The goal and notation

Time runs from noise at $t=0$ to data at $t=1$. Draw a data example $Z\sim p_{\mathrm{data}}$ and independent noise $\varepsilon\sim\mathcal N(0,I_d)$. Lowercase $z$ means a fixed endpoint, while $x$ is a current location. A dot denotes a time derivative; $\nabla_x$ denotes a spatial gradient.

We want a field $u_t(x)$ such that starting with $X_0\sim p_{\mathrm{init}}$ and solving $\dot X_t=u_t(X_t)$ produces the desired data law at the end. Data samples do not directly tell us this field. The solution starts by constructing easier paths, one endpoint at a time.

1. Build conditional paths, then average them

$$\delta_z(A)=\begin{cases}1,&z\in A,\\0,&z\notin A.\end{cases}$$

This defines the Dirac distribution at $z$: a region $A$ has probability 1 if it contains $z$, and 0 otherwise. All probability mass sits at that one point, so drawing $X\sim\delta_z$ always returns $X=z$.

A conditional probability path $p_t(x\mid z)$ describes a noise cloud headed toward one fixed data point. We choose the path to satisfy these endpoint requirements:

$$p_0(\cdot\mid z)=p_{\mathrm{init}},\qquad p_1(\cdot\mid z)=\delta_z.$$

What does $p_1(\cdot\mid z)=\delta_z$ mean? The dot stands for the location variable: this is an equality of entire distributions, not a density evaluated at one location. Given the selected endpoint $z$, the final sample must equal $z$ with probability 1. Every initial noise draw headed toward this endpoint finishes at the same point, rather than leaving a noisy cloud around it. This is a requirement of our chosen bridge, not a property of every possible probability path.

For a Gaussian source, choose differentiable schedules and define:

$$X_t=\alpha_tz+\beta_t\varepsilon,\qquad p_t(x\mid z)=\mathcal N(x;\alpha_tz,\beta_t^2I_d).$$

Use $(\alpha_0,\beta_0)=(0,1)$ and $(\alpha_1,\beta_1)=(1,0)$, with $\beta_t>0$ before the terminal endpoint. The mean moves toward $z$ and the conditional covariance is $\beta_t^2I_d$. Here $\beta_t$ is a standard deviation, not a variance.

Why is it true for this construction? Substitute the terminal schedule values into the sample formula:

$$X_1=\alpha_1z+\beta_1\varepsilon=1\cdot z+0\cdot\varepsilon=z.$$

The noise disappears for every draw of $\varepsilon$. A random variable that always equals $z$ has distribution $\delta_z$, which proves the endpoint condition. For example, if $z=2$, initial noise values $-1$ and $3$ both yield $X_1=2$.

Why combine the individual paths?

So far, we have built a route from noise to one chosen data point. But a generative model must produce the variety of the whole dataset. We therefore need to check that combining these routes gives a distribution that starts as noise and ends with the same frequencies as the data.

Imagine repeating this experiment many times: choose a data point $Z$, draw independent noise $\varepsilon$, and record $X_t=\alpha_tZ+\beta_t\varepsilon$. The cloud of all recorded positions at time $t$ has distribution $p_t$. This is the marginal path: all endpoint choices are included together.

$$p_t(x)=\int p_t(x\mid z)p_{\mathrm{data}}(z)\,dz.$$

Read this as: combine the cloud for each endpoint, weighted by how often that endpoint is selected. The integral performs that weighted combination. This is what the law of total probability means here.

At the start, every recorded position is just noise. Substituting $\alpha_0=0$ and $\beta_0=1$ gives:

$$X_0=0\cdot Z+1\cdot\varepsilon=\varepsilon \quad\Longrightarrow\quad p_0=p_{\mathrm{init}}.$$

The selected data point has no effect yet. Whether we chose $Z=-2$ or $Z=+2$, the starting samples come from the same standard Gaussian.

At the end, every recorded position equals the selected data point. If we select $-2$ half the time and $+2$ half the time, the final positions are also $-2$ half the time and $+2$ half the time. More generally:

$$X_1=1\cdot Z+0\cdot\varepsilon=Z,\qquad Z\sim p_{\mathrm{data}} \quad\Longrightarrow\quad p_1=p_{\mathrm{data}}.$$

This is why we checked the endpoints: the combined path connects exactly the two distributions we want—initial noise and the data distribution. It gives us a target path for learning. A trained generator will later follow a learned field from fresh noise without being given a data endpoint.

For an equally weighted dataset $z_1,\ldots,z_n$, the marginal density before the terminal endpoint is:

$$p_t(x)=\frac1n\sum_{i=1}^n\mathcal N(x;\alpha_tz_i,\beta_t^2I_d).$$

Sampling this mixture only requires choosing an example and adding scaled noise; evaluating its density requires summing the components. The marginal is generally not Gaussian even though every conditional is.

Why does zero noise need a different description? There are two formulas here, and they do different jobs. The sample formula tells us where a sample lands. The density formula describes how probability is spread across locations.

The sample formula still works at the endpoint: $X_1=1\cdot z+0\cdot\varepsilon=z$. There is no division by zero. For $z=2$, every draw returns 2, so the final distribution is perfectly well-defined.

The Gaussian density formula stops working. In one dimension, before the endpoint it is:

$$p_t(x\mid z)=\frac{1}{\sqrt{2\pi}\,\beta_t}\exp\!\left(-\frac{(x-\alpha_tz)^2}{2\beta_t^2}\right),\qquad\beta_t>0.$$

Putting $\beta_1=0$ into this expression divides by zero. As noise shrinks, the Gaussian curve becomes narrower and taller while keeping total area 1. At zero noise, all probability sits at one location. An ordinary density cannot represent that: the area under an ordinary density over a single point is zero, whereas we need probability 1 there.

A point mass describes this final probability directly: for $z=2$, $\Pr(X_1=2)=1$ and $\Pr(X_1\ne2)=0$. We call that distribution $\delta_2$. It does not repair the division by zero or assign a value to the Gaussian density formula. It describes the valid endpoint distribution using point probabilities instead of a density curve.

The same distinction matters for later velocity and score formulas: calling the endpoint a point mass does not make division by $\beta_t=0$ valid. Those formulas must be used before the endpoint or handled with a separately justified limit.

One example for the whole lecture

Let $Z=-2$ or $+2$ with equal probability, and use the straight-line schedules $\alpha_t=t$, $\beta_t=1-t$. The conditional mean is $\alpha_tz=tz$: it is $-2t$ for $z=-2$ and $2t$ for $z=+2$. Combining the two equally likely endpoints gives:

$$p_t(x)=\tfrac12\mathcal N(x;-2t,(1-t)^2)+\tfrac12\mathcal N(x;2t,(1-t)^2).$$

At $t=0$, both components coincide with standard noise. At $t=1/2$, their means are $\pm1$ and their variance is $1/4$. At $t=1$, they become equally weighted atoms at $\pm2$. The marginal mean stays zero, but its variance is $4t^2+(1-t)^2$: shrinking the individual clouds need not shrink the full mixture.

0.35

Two conditional distributions

Orange: fix $z=-2$

$$p_t(x\mid Z=-2)=\mathcal N(x;-2t,(1-t)^2).$$

Green: fix $z=+2$

$$p_t(x\mid Z=+2)=\mathcal N(x;2t,(1-t)^2).$$

Each curve is a full distribution with total probability 1.

Marginal · blue curve

$$\begin{aligned}p_t(x)&=\tfrac12\mathcal N(x;-2t,(1-t)^2)\\&\quad+\tfrac12\mathcal N(x;2t,(1-t)^2).\end{aligned}$$

Combine both endpoints, each with probability $\tfrac12$.

The top panel shows both conditional distributions; the bottom shows their mixture. Each top curve has total probability 1; the bottom gives each endpoint weight 1/2. For $t<1$, curve heights are rescaled per panel to show their shapes, so compare positions and widths rather than peak heights. At $t=1$, labeled spikes show probability masses instead of densities.

2. Differentiate the path to obtain a velocity

Keep the same $(z,\varepsilon)$ while varying time. This supplies a differentiable trajectory; resampling noise independently at each time would give the same snapshots but not this trajectory. Differentiating yields the conditional velocity label:

$$\dot X_t=\dot\alpha_tz+\dot\beta_t\varepsilon.$$

A field must be evaluated at the current location. Solve $x=\alpha_tz+\beta_t\varepsilon$ for $\varepsilon=(x-\alpha_tz)/\beta_t$ and substitute:

$$\boxed{u_t(x\mid z)=\dot\alpha_tz+\frac{\dot\beta_t}{\beta_t}(x-\alpha_tz).}$$

To check the construction, evaluate it on the proposed trajectory:

$$u_t(X_t\mid z)=\dot\alpha_tz+\frac{\dot\beta_t}{\beta_t}(\beta_t\varepsilon)=\dot X_t.$$

Why did we plug $X_t$ into the field? We wanted to check that the field tells each point to move at the speed required by our chosen path. The two speeds match in the equation above. So following this field gives the path $X_t=\alpha_tz+\beta_t\varepsilon$.

Choose many Gaussian noise samples and apply this same formula to each one. Multiplying by $\beta_t$ changes how widely they are spread; adding $\alpha_tz$ moves their center. The result is still Gaussian, now centered at $\alpha_tz$ with standard deviation $\beta_t$ in each direction.

The field does two simple jobs

$$u_t(x\mid z)=\underbrace{\dot\alpha_tz}_{\text{move the center}}+\underbrace{\frac{\dot\beta_t}{\beta_t}(x-\alpha_tz)}_{\text{move toward or away from the center}}.$$

The first term moves every point by the same amount per unit time. This moves the whole cloud's center.

The second term changes the cloud's width. Here $x-\alpha_tz$ tells us where the point is compared with the center. If $\dot\beta_t/\beta_t$ is negative, this term pulls points inward. If it is positive, it pushes them outward.

For example, suppose the center is at 1 and the ratio is $-2$. A point at 1.5 gets an extra velocity $-2(1.5-1)=-1$, pulling it left toward the center. A point at 0.5 gets $-2(0.5-1)=+1$, pulling it right. These are added to the center's own movement. Both points move closer to the center, so the cloud becomes narrower.

Why divide by $\beta_t$? Because $\dot\beta_t$ tells us how fast the width changes, while $\beta_t$ tells us its current size. Their ratio tells the field how strongly to shrink or expand the distances from the center.

Optional math: why the variance check has a factor of 2

In one dimension, variance measures the average squared distance from the center. Call that distance $r_t=X_t-\alpha_tz$. The field above gives $\dot r_t=(\dot\beta_t/\beta_t)r_t$.

Differentiating a square gives $d(r_t^2)/dt=2r_t\dot r_t$. Averaging over the noise draws therefore gives:

$$\frac{d}{dt}\mathbb E[r_t^2]=2\frac{\dot\beta_t}{\beta_t}\mathbb E[r_t^2].$$

Our Gaussian has variance $\mathbb E[r_t^2]=\beta_t^2$. Substituting it gives $2(\dot\beta_t/\beta_t)\beta_t^2=2\beta_t\dot\beta_t$. This matches the derivative of $\beta_t^2$, so the field changes the variance at the correct speed.

In $d$ dimensions, the same calculation gives $2\beta_t\dot\beta_tI_d$, the derivative of the covariance $\beta_t^2I_d$. This is an extra check; you do not need it to understand the two jobs of the field. It applies before $\beta_t$ reaches zero.

Straight-line schedules simplify the label

Here we choose $\alpha_t=t$ and $\beta_t=1-t$. These replace the two coefficients in our general sample formula:

$$X_t=\alpha_tz+\beta_t\varepsilon=tz+(1-t)\varepsilon.$$

The velocity is the time derivative of this position. Keep the endpoint $z$ and the original noise draw $\varepsilon$ fixed while differentiating. Since $d(tz)/dt=z$ and $d((1-t)\varepsilon)/dt=-\varepsilon$,

$$\frac{dX_t}{dt}=z-\varepsilon=u_t(X_t\mid z).$$

The notation $u_t(x\mid z)$ means the velocity assigned to location $x$. Evaluating it at the particle's actual location, $x=X_t$, gives that particle's velocity $dX_t/dt$.

To write the same velocity using the current location instead of the original noise, solve $x=tz+(1-t)\varepsilon$ for $\varepsilon$ and substitute:

$$u_t(x\mid z)=z-\frac{x-tz}{1-t}=\frac{z-x}{1-t},\qquad t<1.$$

For $z=2$, $\varepsilon=-1$, $t=1/4$, the input is $x=-1/4$. Both velocity formulas give $3$. The noise-based label remains finite, but this does not remove the location-based field's endpoint singularity: different incoming trajectories collapse to the same $z$.

The schedule determines the label. For $\alpha_t=\sin(\pi t/2)$ and $\beta_t=\cos(\pi t/2)$, differentiation instead gives $\frac\pi2[\cos(\pi t/2)z-\sin(\pi t/2)\varepsilon]$. Reusing $z-\varepsilon$ would target the wrong dynamics. No ODE solve or density evaluation is needed to construct either label.

3. Combine the two possible velocities

Start with the two data values

$-2$ and $+2$ are the two data values we chose for this example. They are fixed destinations, not values calculated by the model. One group of paths ends at $-2$; the other ends at $+2$. We choose the two destinations equally often.

At time $t$, the first group's Gaussian cloud is centered at $-2t$ and the second at $+2t$. For example, halfway through, at $t=0.5$, their centers are $-1$ and $+1$. At the end, those centers reach the data values $-2$ and $+2$.

Now imagine seeing a point at location $x$, without knowing which group it came from. Each group suggests a velocity for that point. We need to combine these suggestions into one velocity that uses only the current location and time.

$$u_t(x)=\mathbb E\!\left[u_t(X_t\mid Z)\mid X_t=x\right].$$

This says: at the current location $x$, average the velocities suggested by the possible endpoints. The previous section gave a velocity when the endpoint $z$ is known. Now we need one velocity when that endpoint is unknown. The expectation $\mathbb E$ means a probability-weighted average; it is the rule we will expand below. Section 4 proves that this rule moves the whole distribution correctly.

Which two velocities are we averaging?

Our example has only two possible data endpoints, $z=-2$ and $z=+2$. Substitute each one into the formula from the previous section, $u_t(x\mid z)=(z-x)/(1-t)$:

$$u_t(x\mid-2)=\frac{-2-x}{1-t},\qquad u_t(x\mid+2)=\frac{2-x}{1-t}.$$

These are two different velocity suggestions for the same current location $x$. The first assumes the endpoint is $-2$; the second assumes it is $+2$.

Where $w_-$ and $w_+$ come from

Give short names to the probabilities of those two endpoints after seeing the current location:

$$w_-=\Pr(Z=-2\mid X_t=x),\qquad w_+=\Pr(Z=+2\mid X_t=x).$$

The minus and plus signs are just labels for the two endpoints. Neither weight is negative. They depend on $t$ and $x$, and add to 1 because these are the only two possibilities.

A probability-weighted average means “first value × its probability + second value × its probability.” Applying that definition to the two velocity suggestions gives:

$$\boxed{u_t(x)=w_-\,u_t(x\mid-2)+w_+\,u_t(x\mid+2).}$$

$w_-\,u_t$ means multiplication: multiply a velocity by the fraction assigned to it. It is not a new symbol called “$w_u$.” This two-term sum is simply the expectation above written out for our two endpoints. With many possible endpoints, we sum over them; for a continuous range of endpoints, that sum becomes an integral.

To calculate the weights, use Bayes' rule. In this example both endpoints have the same starting probability, $1/2$, which cancels from the numerator and denominator:

$$w_+=\frac{p_t(x\mid+2)}{p_t(x\mid-2)+p_t(x\mid+2)},\qquad w_-=1-w_+.$$

Read the fraction as “height of the $+2$ curve at $x$, divided by the sum of both curve heights there.” A larger share of the density gives that endpoint's velocity a larger weight. The figure below shows this without requiring any calculation by hand.

What is a weight?

A weight is the fraction of a group's velocity we use. For example, weights of 80% and 20% mean “use 80% of the first velocity plus 20% of the second.” The two weights always add to 100%.

How do we choose them? Look at the height of each Gaussian curve at the same location $x$. A taller curve means that group produces more samples near this location. We therefore give its velocity a larger share. With equally common groups, equal curve heights give weights of 50% each.

So the initial 50–50 choice describes how often we pick each destination overall. The weights describe how likely each group is for the particular location we are now looking at. These are different questions.

How to read and move this figure

Start with the defaults: $t=0.5$ and $x=0$. The line is halfway between the two clouds, their heights match, and each velocity gets 50%. Their opposite directions cancel. Keep time fixed and move $x$ right: the right-hand cloud becomes more likely at that location, so its percentage increases.

Share of the orange velocity ($z=-2$)
Share of the green velocity ($z=+2$)
$u_t(x\mid z=-2)$
$u_t(x\mid z=+2)$
combined velocity (blue)
Orange: paths ending at $-2$. Green: paths ending at $+2$. Both curve heights are compared at the dashed line. The blue arrow is the resulting velocity; long arrows are clipped to fit the plot, while the numbers show their full values.

The same rule for a continuous range of endpoints

$$u_t(x)=\int u_t(x\mid z)p_t(z\mid x)\,dz,\qquad p_t(z\mid x)=\frac{p_t(x\mid z)p_{\mathrm{data}}(z)}{p_t(x)}.$$

The integral does the same job as our two-term sum: combine the velocity suggestions using the probabilities of their endpoints. Here $p_t(z\mid x)$ is a density over endpoints, so probabilities come from integrating it over a region. This gives one field of $t$ and $x$ without requiring the sampler to know a future data point.

At exactly $x=0$ in this symmetric example, the combined velocity is zero. That single starting point has probability zero under continuous Gaussian noise, so its staying at zero does not change the final distribution.

4. Check that the velocity moves the distribution correctly

Why do we need another equation? We have proposed a velocity $u_t(x)$. We must now check that moving samples with this velocity produces the distribution $p_t$ we chose. The continuity equation is the test: it connects how samples move to how their density changes.

First, read the derivative notation

$$\partial_t p_t(x)\;=\;\frac{\partial p_t(x)}{\partial t}.$$

$\partial_t$ is shorthand for a derivative with respect to time. The denominator has not disappeared; it is hidden in this shorter notation. The symbol $\partial$ is called “partial,” not delta. We use it because $p_t(x)$ depends on both time $t$ and location $x$.

Here we differentiate the density $p_t(x)$ at a fixed location $x$ with respect to time; $dX_t/dt$ instead differentiates a sample's position $X_t$ to give its velocity. Think of watching traffic get denser at one spot versus following one car to measure its speed.

Probability can enter or leave a region

Imagine a small interval on a number line. Samples move into it through one boundary and out through the other. If more probability enters than leaves, the probability inside increases. If more leaves than enters, it decreases. Moving samples does not create or destroy probability.

To measure the flow across a boundary, we need both the density there and how fast the samples move. Their product is called probability flux:

$$J_t(x)=p_t(x)u_t(x).$$

In one dimension, $J_t(x)$ measures the signed rate of probability flowing across location $x$: positive for rightward flow, negative for leftward flow. For an interval $[a,b]$, the amount of probability inside is $\int_a^b p_t(x)\,dx$. Its change is:

$$\frac d{dt}\int_a^b p_t(x)\,dx=\underbrace{J_t(a)}_{\text{flow through left boundary}}-\underbrace{J_t(b)}_{\text{flow through right boundary}}.$$

This is the bookkeeping rule “change inside = inflow − outflow.” It is the reason for the continuity equation, rather than a new assumption about the data.

Write the same rule at each location

Take the interval above to be $[x,x+h]$, where $h>0$ is a small width. The probability inside is approximately $p_t(x)h$: density times width.

The flow rule measures the change in that probability per unit time, so we differentiate $p_t(x)h$. The interval stays fixed: $h$ does not change with time, so it comes outside the derivative:

$$\frac{\partial}{\partial t}\big[p_t(x)h\big]=h\frac{\partial p_t(x)}{\partial t}.$$

That is why the left side below is $h\,\partial p_t(x)/\partial t$, rather than just $p_t(x)h$. Applying “inflow − outflow” gives:

$$h\frac{\partial p_t(x)}{\partial t}\approx J_t(x)-J_t(x+h).$$

Divide by $h$ to get the change in density, rather than the change in probability across the whole interval:

$$\frac{\partial p_t(x)}{\partial t}\approx-\frac{J_t(x+h)-J_t(x)}{h}.$$

As $h$ shrinks to zero, the fraction on the right becomes $\partial J_t(x)/\partial x$ by the definition of a derivative. This gives the first equality below; substituting our definition $J_t(x)=p_t(x)u_t(x)$ gives the second:

$$\frac{\partial p_t(x)}{\partial t}=-\frac{\partial J_t(x)}{\partial x}=-\frac{\partial}{\partial x}\big[p_t(x)u_t(x)\big].$$

The right side compares the flow at nearby locations. If flow increases as we go from left to right, more probability leaves a small interval than enters it, so its density falls. That explains the minus sign. If flow decreases, probability builds up instead.

For images and other vectors, there are many coordinate directions. Divergence, written $\operatorname{div}$, adds up the changes in flow across those directions. It replaces the one-dimensional derivative $\partial/\partial x$. The same equation is then written:

$$\boxed{\partial_t p_t=-\operatorname{div}(p_tu_t).}$$

Read it as: the density's rate of change equals minus the net outward flow of probability. This is called the continuity equation. The product inside the brackets matters: we track the flow of probability, so we need density times velocity, not velocity alone.

Optional: the algebra behind the local form

The fundamental theorem of calculus gives $J_t(a)-J_t(b)=-\int_a^b\partial_xJ_t(x)\,dx$. Substitute this into the interval equation. Since it holds for every interval, the quantities inside the integrals agree, giving $\partial_t p_t=-\partial_xJ_t$ for smooth densities and fields.

In $d$ dimensions, $\operatorname{div}J_t=\sum_{j=1}^d\partial J_{t,j}/\partial x_j$: differentiate each flow component along its own coordinate, then add. Expanding the product gives $\partial_t p_t=-u_t\cdot\nabla p_t-p_t\operatorname{div}u_t$. Along an ODE trajectory where the density is positive, this also gives $d\log p_t(X_t)/dt=-\operatorname{div}u_t(X_t)$.

What the next proof must show: our averaged velocity satisfies this equation for the chosen marginal density $p_t$. That will establish that the field moves the full distribution along the intended path.

The marginalization proof

The divergence comes from substituting the continuity equation we just derived. For a fixed destination $z$, the conditional density and velocity obey the same rule:

$$\partial_t p_t(x\mid z)=-\operatorname{div}_x\!\big[p_t(x\mid z)u_t(x\mid z)\big].$$

So whenever we see $\partial_t p_t(x\mid z)$ inside the integral, we replace it with the right side above. The subscript $x$ means divergence measures changes in flow across locations; in one dimension it is just $\partial/\partial x$.

$$\begin{aligned} \int\partial_t p_t(x\mid z)p_{\mathrm{data}}(z)\,dz &=-\int\operatorname{div}_x\!\big[p_t(x\mid z)u_t(x\mid z)\big]p_{\mathrm{data}}(z)\,dz\\ &=-\operatorname{div}_x\!\left(\int p_t(x\mid z)u_t(x\mid z)p_{\mathrm{data}}(z)\,dz\right). \end{aligned}$$

Why can divergence move outside? We are averaging over destinations $z$, but differentiating with respect to location $x$. Differentiation is linear, and the averaging weights $p_{\mathrm{data}}(z)$ do not depend on $x$, so we can average the flows first and then take their divergence, assuming the regularity conditions stated below.

Now apply this substitution to the marginal density $p_t(x)=\int p_t(x\mid z)p_{\mathrm{data}}(z)\,dz$ and finish the proof:

$$\begin{aligned} \partial_t p_t(x) &=\int\partial_t p_t(x\mid z)p_{\mathrm{data}}(z)\,dz\\ &=-\operatorname{div}\!\left(\int p_t(x\mid z)u_t(x\mid z)p_{\mathrm{data}}(z)\,dz\right)\\ &=-\operatorname{div}\!\left(p_t(x)\int u_t(x\mid z)\frac{p_t(x\mid z)p_{\mathrm{data}}(z)}{p_t(x)}\,dz\right)\\ &=-\operatorname{div}(p_tu_t)(x). \end{aligned}$$

The third line multiplies and divides by $p_t(x)$, exposing the posterior weights. Thus the marginal field has precisely the probability flux needed to follow the marginal path.

This argument assumes derivatives can pass through the integral and the dynamics have the required well-posedness and uniqueness. Start with the correct initial law and interpret singular terminal distributions as limits. The construction supplies a valid field, not necessarily the only one: adding $v_t$ with $\operatorname{div}(p_tv_t)=0$ leaves the same continuity equation.

5. Derive the conditional and marginal scores

The score $s_t(x)=\nabla_x\log p_t(x)$ points toward increasing log density. It is a spatial derivative, not a velocity or probability value. For the Gaussian conditional, terms independent of $x$ can be grouped into $C$:

$$\log p_t(x\mid z)=C-\frac{\|x-\alpha_tz\|^2}{2\beta_t^2}.$$

Differentiate with respect to $x$, keeping $t$ and $z$ fixed. The constant $C$ has derivative zero, and $\alpha_tz$ and $\beta_t$ are constant with respect to $x$. In one dimension, differentiating $(x-\alpha_tz)^2$ gives $2(x-\alpha_tz)$; for a vector, the gradient does this for each coordinate:

$$\begin{aligned} s_t(x\mid z) &=\nabla_x\log p_t(x\mid z)\\ &=0-\frac{1}{2\beta_t^2}\nabla_x\|x-\alpha_tz\|^2\\ &=-\frac{1}{2\beta_t^2}\,2(x-\alpha_tz)\\ &=-\frac{x-\alpha_tz}{\beta_t^2}. \end{aligned}$$

The score points back toward the conditional mean. For a sampled point $X_t=\alpha_tz+\beta_t\varepsilon$, subtracting the mean leaves $X_t-\alpha_tz=\beta_t\varepsilon$. Substitute this into the score formula:

$$s_t(X_t\mid z)=-\frac{X_t-\alpha_tz}{\beta_t^2}=-\frac{\beta_t\varepsilon}{\beta_t^2}=-\frac{\varepsilon}{\beta_t},\qquad \beta_t>0.$$

Why does smaller noise make the score steeper? In one dimension, the score's slope with respect to position is $\partial s_t(x\mid z)/\partial x=-1/\beta_t^2$. A smaller $\beta_t$ makes this slope larger in magnitude: the same step away from the mean produces a stronger score pointing back toward it. This describes the narrower Gaussian before the zero-noise endpoint.

Scores obey the same posterior averaging rule

$$\begin{aligned} s_t(x)&=\frac{\nabla_xp_t(x)}{p_t(x)}\\ &=\int\frac{\nabla_xp_t(x\mid z)}{p_t(x\mid z)} \frac{p_t(x\mid z)p_{\mathrm{data}}(z)}{p_t(x)}\,dz\\ &=\int s_t(x\mid z)p_t(z\mid x)\,dz. \end{aligned}$$

Differentiate the mixture and insert Bayes' rule; this does not require the mixture itself to be Gaussian. Densities are averaged with prior data weights, while velocities and scores use posterior endpoint weights.

A score and a velocity answer different questions

The velocity tells a sample how to move to follow the chosen path. The score tells us which direction increases the density at the current time. They need not point the same way. Also, a zero score does not necessarily mean a density maximum: at the symmetric midpoint between two modes, the slope can be zero even in the valley.

6. Add random movement while keeping the same distributions

What are we trying to do? Our ODE already moves a cloud of samples through the desired densities $p_t$. We now want to let individual samples wander randomly while keeping the whole cloud's density the same at each time.

Start with the ODE. Its small position change is velocity times elapsed time:

$$\frac{dX_t}{dt}=u_t(X_t)\qquad\Longleftrightarrow\qquad dX_t=u_t(X_t)\,dt.$$

To add random movement, use $\sigma_t\,dW_t$. Here $W_t$ is Brownian motion, $dW_t$ is its fresh random increment, and the chosen number $\sigma_t\ge0$ controls its size. Over a short interval $h$, this random displacement is distributed as $\sigma_t\sqrt h\,\xi$ when the coefficient is held fixed, with $\xi\sim\mathcal N(0,I_d)$. We let $\sigma_t$ depend only on time.

Random kicks spread the cloud beyond the path we chose. Introduce an extra velocity $c_t(x)$ to compensate; we still need to work out what it should be:

$$dX_t=[u_t(X_t)+c_t(X_t)]\,dt+\sigma_t\,dW_t.$$

Find the correction from the density change

In one dimension, the continuity equation gave the density change from directed movement. With Brownian noise, it gains a spreading term. This extended rule is called the Fokker–Planck equation:

$$\frac{\partial p_t}{\partial t}= \underbrace{-\frac{\partial}{\partial x}(p_tu_t)}_{\text{original density change}} \quad\underbrace{-\frac{\partial}{\partial x}(p_tc_t)}_{\text{effect of correction}} \quad+\underbrace{\frac{\sigma_t^2}{2}\frac{\partial^2p_t}{\partial x^2}}_{\text{effect of noise}}.$$

$\partial^2p_t/\partial x^2$ means differentiate the density twice with respect to position. The first term already produces the desired path. We need the last two terms to cancel.

$p_t$ is already fixed, but $c_t$ is still unknown. We introduced $c_t$ as the extra velocity we want to design. The product $p_tc_t$ is the probability flow caused by that extra velocity; it is not fixed yet because we have not chosen $c_t$. We are solving for this new velocity while keeping the desired density $p_t$ unchanged.

One way to satisfy the cancellation requirement is to define $c_t$ so that $p_tc_t=(\sigma_t^2/2)\,\partial p_t/\partial x$. The right side is determined by our chosen density and noise scale; dividing it by $p_t$ determines $c_t$ wherever $p_t>0$. First, check that this choice works by substituting the product into the correction term:

$$\begin{aligned} -\frac{\partial}{\partial x}(p_tc_t) &=-\frac{\partial}{\partial x}\left[\frac{\sigma_t^2}{2}\frac{\partial p_t}{\partial x}\right]\\ &=-\frac{\sigma_t^2}{2}\frac{\partial^2p_t}{\partial x^2}. \end{aligned}$$

$\sigma_t$ depends only on time, so $\sigma_t^2/2$ is constant when differentiating with respect to $x$. Differentiating $\partial p_t/\partial x$ once more gives the second derivative. The minus sign was already outside the correction term. Now add the noise term:

$$\underbrace{-\frac{\sigma_t^2}{2}\frac{\partial^2p_t}{\partial x^2}}_{\text{correction}}+\underbrace{\frac{\sigma_t^2}{2}\frac{\partial^2p_t}{\partial x^2}}_{\text{noise}}=0.$$

They are the same quantity with opposite signs, so they cancel whether the second derivative itself is positive or negative. Finally, divide our choice for $p_tc_t$ by $p_t$ to get the extra velocity:

$$c_t(x)=\frac{\sigma_t^2}{2}\frac{\partial_xp_t(x)}{p_t(x)}=\frac{\sigma_t^2}{2}\,\partial_x\log p_t(x)=\frac{\sigma_t^2}{2}s_t(x).$$

The middle equality is the derivative rule for a logarithm, and the last is our definition of the score. In several dimensions, replace the spatial derivative by the gradient $\nabla_x$; the same score correction results.

The total directed velocity, called the drift, is $b_t=u_t+(\sigma_t^2/2)s_t$. The score correction counters the additional spreading from the kicks, leaving the original density evolution. Setting $\sigma_t=0$ removes both added terms and gives back the ODE.

$\sigma_t$ controls fresh noise added during this movement; $\beta_t$ specifies the noise scale of the probability path we chose earlier. They are different quantities.

Putting the drift and the random movement together gives the final SDE:

$$\boxed{dX_t=\left[u_t(X_t)+\frac{\sigma_t^2}{2}s_t(X_t)\right]dt+\sigma_t\,dW_t.}$$
Same distributions, different trajectories. The ODE fixes a path once its initial sample is chosen; the SDE receives fresh Brownian increments. The theorem requires the correct initial law, exact fields, and sufficient regularity. State-dependent diffusion needs additional terms. Learned fields, finite solver steps, and singular endpoints require separate numerical care.

The six formulas to keep handy

Conditional means we fix one endpoint $z$. Marginal means we combine all possible endpoints. These formulas use $X_t=\alpha_tZ+\beta_t\varepsilon$ and apply where $\beta_t>0$.

QuantityConditional: one endpointMarginal: all endpoints
Probability density
Where are the samples?
$$p_t(x\mid z)=\mathcal N(x;\alpha_tz,\beta_t^2I_d).$$$$p_t(x)=\int p_t(x\mid z)p_{\mathrm{data}}(z)\,dz.$$
Velocity
How do samples move?
$$\begin{aligned}u_t(x\mid z)&=\dot\alpha_tz\\&\quad+\frac{\dot\beta_t}{\beta_t}(x-\alpha_tz).\end{aligned}$$$$u_t(x)=\int u_t(x\mid z)p_t(z\mid x)\,dz.$$
Score
Which way does density increase?
$$s_t(x\mid z)=-\frac{x-\alpha_tz}{\beta_t^2}.$$$$s_t(x)=\int s_t(x\mid z)p_t(z\mid x)\,dz.$$

The density row uses how often each endpoint occurs in the data, $p_{\mathrm{data}}(z)$. The velocity and score rows use how likely that endpoint is after seeing the current location $x$: $p_t(z\mid x)=p_t(x\mid z)p_{\mathrm{data}}(z)/p_t(x)$. That is why their weights differ.

For the straight-line path, substitute $\alpha_t=t$, $\beta_t=1-t$, $\dot\alpha_t=1$, and $\dot\beta_t=-1$. Along a conditional trajectory, the velocity in the table is $u_t(X_t\mid z)=dX_t/dt$.

Sources and reading route

Based on my nine-page handwritten Lecture 2 — constructing a training target (13 December 2025): probability paths on pages 1–4, conditional velocities on pages 4–5, marginalization and continuity on pages 5–7, and scores and Fokker–Planck on pages 7–9. The numerical examples expand the source notes.

Companion material: MIT 6.S184 Lecture 2 slides and the 2025 course notes and recordings. Review Lecture 1's ODE/SDE foundations, then continue to Lecture 3's training objectives.