← All Posts
Deep Learning · Diffusion & Flow Models · Large Language Diffusion Models

Quantizing a Masked Diffusion LM: Weights, Trajectories, and Calibration

Before you start: Masked diffusion training, sampling, and GPTQ fundamentals.

A four-bit weight file is smaller than a sixteen-bit file. That does not tell you how much live GPU memory the sampler uses, whether its kernels run faster, or whether small numerical errors change which tokens are revealed. A masked diffusion language model needs evaluation at the level of both individual denoising passes and the complete trajectory.

01Compressed weightsBits, groups, scales, zero points
02Denoising passActivations depend on the mask state
03Token decisionLogits affect confidence and selection
04TrajectoryChanged tokens alter later context
A local quantization error can change the sampler’s future inputs. Evaluate the full policy, not only one layer’s reconstruction error.

Start with a uniform quantizer

$$q=\operatorname{clip}\!\left(\operatorname{round}(w/s)+z,q_{\min},q_{\max}\right),\qquad \widehat w=s(q-z).$$

The integer q stores a quantized level; s is scale and z is zero point. Symmetric signed quantization often uses z=0. Group-wise quantization gives small groups separate scales so one outlier does not set the resolution for an entire matrix. Metadata and unquantized layers add to the stored size, so parameter count times bit width is only the main payload estimate.

A fixed-scale illustration with rounding and no clipping for these values. The interactive control uses a separately declared range and level count.
WeightWith symmetric scale 0.25Error
−0.90−1.00−0.10
−0.20−0.25−0.05
0.100.00−0.10
0.550.50−0.05
0.951.00+0.05
Change the number of quantization levels for five synthetic weights. The figure computes rounding error directly; it does not simulate a trained language model.

Small weight error is not the whole objective

For a linear layer, the output error is $(W-\widehat W)x$. Its size depends on the activations x, not only the weight error. Calibration methods use representative activations to choose or compensate for quantization. GPTQ and AWQ make different choices about reconstruction and salient channels; the existing quantization notes explain their base algorithms.

In a masked diffusion model, x changes with mask fraction, revealed context, sequence length, and sampler decisions. Calibrating only on clean text can miss states encountered during generation. A calibration set should cover the actual prompt/response structure and a range of denoising states, including states produced by the intended sampler when feasible.

Design the calibration experiment

Freeze the checkpoint, tokenizer, prompt template, output length, and sampler. Collect activation examples across masking levels and representative tasks. Keep calibration data separate from final evaluation. Compare clean-only calibration, randomly masked states, and trajectory-derived states under the same calibration budget. This tests coverage rather than assuming one source is best.

If calibration uses full-precision trajectories, quantized sampling can later visit different states. Measure that shift. A second calibration round using quantized trajectories is an experimental option, but it must not use held-out test answers to improve the model. Record masking and position-selection policies with the calibration artifact.

Separate local error from decision changes

Use identical frozen states for local comparisons, then allow trajectories to evolve independently for end-to-end evaluation.
LevelMeasurementQuestion
LayerRelative output error on held-out activationsWhich operations are numerically sensitive?
LogitsDistribution divergence or top-token agreementAre predictions changing at fixed states?
SelectionAgreement on positions chosen for reveal/remaskingDid confidence ordering change?
TrajectoryValidity, task accuracy, diversity, terminationWhat happens after decisions feed back?
SystemLive memory, peak memory, latency, throughputDoes the deployed implementation deliver a useful operating point?

A tiny logit perturbation can swap two nearly tied confidence values and change the reveal order. That creates a discrete difference even when average reconstruction error is small. Conversely, a visible numerical error may not change any important decision. Inspect sensitivity near decision boundaries instead of assuming every layer error has equal task impact.

Compressed storage, live memory, and speed are separate results

Weight-only quantization may dequantize into higher precision inside a kernel. Activations, temporary buffers, logits, model copies, and allocator behavior can dominate live memory. Report artifact size, resident memory after loading, and peak memory during generation separately. Measure after warm-up and synchronize asynchronous device work.

Low-bit storage does not guarantee accelerated arithmetic on the target hardware. Kernel support, fusion, group size, batch shape, and dequantization overhead affect latency. Compare configurations at matched task quality or show the full quality–latency–memory frontier. Record actual denoiser calls and tokens processed, especially if a changed sampler uses more iterations to recover quality.

What transfers to another model?

A quantization recipe, a calibration dataset, and already quantized parameters are three different objects. Test recipe transfer across checkpoints, architectures, and tasks while keeping the comparison controlled. Transfer of one successful configuration does not imply that every sensitive layer, group size, or calibration distribution remains appropriate.

Publish an experiment manifest with precision per module, calibration source and mask coverage, sampler configuration, hardware, kernels, task metrics, uncertainty, and failure examples. This makes “what compresses” a reproducible question and “what breaks” a diagnosable one. The next chapter examines another trajectory-changing intervention: starting diffusion from an autoregressive draft.

Check your understanding: Why can a low average layer error still change the final answer?

The error can change a token or position-selection decision near a tie. That decision changes later context, so the remaining denoising trajectory can diverge.

Sources and further reading