Introduction to Flow & Diffusion Models
A generative model turns easy-to-sample noise into data-like objects. Flow models move samples deterministically after initialization; diffusion models add random motion along the way. This lecture explains the dynamics and how to simulate them, assuming a useful field is already available. Lectures 2 and 3 construct its targets and explain how to learn it.
1. Generation is a distribution problem
Represent an object as a vector $x\in\mathbb R^d$: an RGB image has $d=3HW$ pixel values, a video adds a frame dimension, and molecular coordinates or audio require their own representations. We observe a finite dataset from an unknown distribution $p_{\mathrm{data}}$. The objective is to generate new samples with that distribution's structure, rather than simply replay stored examples.
For conditional generation, fix side information $y$—a prompt, class, or another image—and target $p_{\mathrm{data}}(x\mid y)$. A realistic cat image can still be wrong for a dog prompt. Distribution matching formalizes the task, but density alone is not a complete measure of human quality or prompt alignment.
Start with a convenient source such as $X_0\sim p_{\mathrm{init}}=\mathcal N(0,I_d)$. Let $X_t$ have distribution $p_t$ as it moves. The ideal endpoints are:
We use $t=0$ for noise and $t=1$ for data throughout this series. A probability path $(p_t)$ describes a sequence of distribution snapshots. A trajectory describes one sample's motion through those snapshots. A learned model and numerical solver generally match the desired endpoint only approximately.
2. Trajectory, vector field, ODE, and flow map
These are four views of the same deterministic motion:
- Trajectory $X_t$: the position of one particle as time varies.
- Vector field $u_t(x)$: the velocity prescribed at every location and time. Its output has the same dimension as $x$.
- ODE $\dot X_t=u_t(X_t)$: the particle's velocity must match the field at its current position.
- Flow map $\psi_t(x_0)$: the position reached at time $t$ when the ODE starts at $x_0$.
The field supplies local arrows; the ODE follows them; the flow map records the resulting displacement of every possible starting point. The map is not the velocity itself.
Why a fixed starting point has one future
Spatial Lipschitz regularity, suitable time regularity, and growth bounds give existence and uniqueness over the interval of interest. For example, a continuously differentiable field with bounded spatial derivatives is spatially Lipschitz. Local regularity alone does not rule out finite-time blow-up.
Under the corresponding forward and backward well-posedness, trajectories cannot meet at the same point and time after starting separately: reversing from that meeting would violate uniqueness. The flow is then invertible by solving backward. A limiting collapse into point masses can violate endpoint regularity; Lecture 2 handles such endpoints as distributional limits.
A solvable example: linear contraction
Use $u(x)=-\lambda x$ with $\lambda>0$. We write $\lambda$ for this contraction rate, keeping $\theta$ for neural-network parameters later. The solution is:
Check both requirements: at $t=0$ this equals $x_0$, and differentiating gives $-\lambda e^{-\lambda t}x_0=-\lambda X_t$. Positive points move left and negative points move right; their distances from zero shrink exponentially without reaching zero in finite time.
3. A deterministic map can produce random samples
Fixing $X_0=x_0$ gives one trajectory. Drawing $X_0\sim p_{\mathrm{init}}$ makes the output random even though the same deterministic map is applied to every draw:
Pushforward is the name for “sample from the initial distribution, apply the map, and look at the resulting distribution.” The symbol $\#$ refers to transporting a distribution, not an extra operation on one particle.
Formally, for a region $A$, $\Pr(X_t\in A)=\Pr(X_0\in\psi_t^{-1}(A))$. Probability in an output region equals the probability that started in its preimage. No mass is created or destroyed.
For the contraction example with $X_0\sim\mathcal N(0,1)$:
Halving every sample halves the standard deviation and quarters the variance. The particle-level contraction and the narrowing Gaussian cloud are the same phenomenon. A generative flow seeks $(\psi_1)_\#p_{\mathrm{init}}\approx p_{\mathrm{data}}$; it matches distributions rather than assigning every noise draw a predetermined training example.
4. Turn velocity queries into samples with Euler's method
A neural vector field usually has no closed-form flow map. Approximate its ODE on a grid $t_k=kh$, with $h=1/n$. The derivative relation gives:
Each step follows the current arrow briefly, then queries a new arrow at the updated position. The discrete state $X_k$ approximates the continuous state at time $t_k$.
Accuracy and stability are different questions
For $u(x)=-\lambda x$, Euler gives $X_{k+1}=(1-\lambda h)X_k$. Over the unit interval:
With $\lambda=1$, $X_0=4$, and two steps of $h=0.5$, the states are $4\to2\to1$. The exact endpoint is $4e^{-1}\approx1.472$. More sufficiently small steps reduce this approximation error.
For sufficiently smooth dynamics, Taylor expansion gives a one-step local error $O(h^2)$; accumulated over a fixed interval, Euler's global error is $O(h)$ under the usual stability and regularity assumptions. Halving a sufficiently small step approximately halves the error and doubles field evaluations.
A stable differential equation can still have an unstable numerical approximation. Here Euler contracts only when:
For $1<\lambda h<2$, it alternates signs while decaying; above 2 it grows in magnitude. The exact solution never oscillates. Smaller stable steps or a better solver address numerical error, but cannot repair a wrongly learned field.
A neural flow sampler
A network $u_t^\theta(x)$ receives the current sample and time and returns a velocity of the same shape. After training, hold $\theta$ fixed:
- Draw fresh $X_0\sim\mathcal N(0,I_d)$ and choose $h=1/n$.
- For $k=0,\ldots,n-1$, compute $X_{k+1}=X_k+h\,u_{t_k}^\theta(X_k)$.
- Return $X_n$ as the approximate generated sample.
A batch follows the same field from different initial draws. With a deterministic field and solver, the same initial draw gives the same output; diversity comes from the initial noise.
5. Brownian motion adds randomness through time
A stochastic process $(X_t)$ is a collection of random variables indexed by time. One realization is a sample path. Brownian motion $W_t\in\mathbb R^d$ has:
- $W_0=0$.
- Independent increments on disjoint intervals, with $W_t-W_s\sim\mathcal N(0,(t-s)I_d)$ for $s<t$.
- Continuous sample paths with probability one.
Its paths are almost surely nowhere differentiable. Thus $dW_t$ is stochastic-increment notation; it is not an ordinary velocity multiplied by $dt$.
Drift plus diffusion
The drift $b_t$ supplies directed motion; the nonnegative, time-only diffusion coefficient $\sigma_t$ scales random increments. Using $b_t$ distinguishes an SDE drift from an ODE velocity $u_t$: they need not be equal when both processes are meant to follow the same distributions.
With zero diffusion this is an ODE. With nonzero diffusion, fixing the initial point still permits different paths because of fresh random increments. Under appropriate Lipschitz and growth conditions, a unique strong solution means the path is determined once both the initial condition and Brownian realization are fixed.
Why the noise increment scales as a square root
Over an interval of length $h$, Brownian motion has variance $hI_d$. Therefore:
Scaling by $\sqrt h$ multiplies variance by $h$. Scaling by $h$ would give variance $h^2$; over $1/h$ independent increments, that erroneous accumulated variance tends to zero rather than one.
The Euler–Maruyama sampler
Draw a fresh independent standard Gaussian $\varepsilon_k$ at each step, also independent of the initial sample. For a learned diffusion model, use its learned drift $b_t^\theta$, initialize from the chosen source distribution, and return the final state just as for the flow sampler. With $h=0.01$ and $\sigma=0.5$, one noise increment has standard deviation $0.05$ and variance $0.0025$.
6. One example connects deterministic and stochastic motion
Add noise to the same linear contraction, with constants $\lambda>0$ and $\sigma\ge0$:
This is the Ornstein–Uhlenbeck process. Starting from fixed $X_0=x_0$, multiplying by the integrating factor $e^{\lambda t}$ gives:
The stochastic integral has mean zero. Its variance is the integral of the squared deterministic coefficient, so:
The mean follows the ODE, but paths spread around it. In the long run, inward drift balances diffusion at variance $\sigma^2/(2\lambda)$. This formula assumes a fixed starting value; an independent random $X_0$ contributes the additional variance $e^{-2\lambda t}\operatorname{Var}(X_0)$.
The figure uses Euler–Maruyama, so it approximates the continuous process. The displayed long-run variance is the continuous-time value. Setting $\sigma=0$ collapses the paths to the same deterministic numerical trajectory.
What remains to build a trained generator?
We now know how to generate samples given useful dynamics. Adding arbitrary noise to an ODE does not automatically preserve its target distribution. Lecture 2 constructs probability paths, derives their target fields, and proves the score correction that makes an SDE follow the same marginals. Lecture 3 turns conditional targets into flow-matching and score-matching objectives.
Short exercises with solutions
Why can a deterministic flow generate different images?
The initial noise is random. Each draw follows a deterministic map, so different starting draws can produce different outputs. Deterministic dynamics do not mean a constant output.
What happens to a standard Gaussian under $\psi(x)=x/2$?
The output is $\mathcal N(0,1/4)$. Its standard deviation is halved, and its variance is quartered. This is a pushforward of the entire distribution.
Give an oscillating Euler step for $\dot X=-3X$.
Choose $h=0.5$. The multiplier is $1-3h=-0.5$, so signs alternate while magnitudes shrink. With $h=1$, the multiplier is $-2$ and the approximation diverges.
Why must Euler–Maruyama draw fresh noise at every step?
Brownian increments on disjoint intervals are independent. Reusing the same Gaussian draw couples those increments and changes the process being simulated.
Sources and reading route
Consolidated from my Lecture 1 notes, Generative AI with SDEs, with the contraction, numerical-error, and variance calculations worked out here. Companion material: MIT 6.S184 Lecture 1 slides and the 2025 course notes, recordings, and labs.