Quantizing a Masked Diffusion LM: Weights, Trajectories, and Calibration
Before you start: Masked diffusion training, sampling, and GPTQ fundamentals.
A four-bit weight file is smaller than a sixteen-bit file. That does not tell you how much live GPU memory the sampler uses, whether its kernels run faster, or whether small numerical errors change which tokens are revealed. A masked diffusion language model needs evaluation at the level of both individual denoising passes and the complete trajectory.
Start with a uniform quantizer
The integer q stores a quantized level; s is scale and z is zero point. Symmetric signed quantization often uses z=0. Group-wise quantization gives small groups separate scales so one outlier does not set the resolution for an entire matrix. Metadata and unquantized layers add to the stored size, so parameter count times bit width is only the main payload estimate.
| Weight | With symmetric scale 0.25 | Error |
|---|---|---|
| −0.90 | −1.00 | −0.10 |
| −0.20 | −0.25 | −0.05 |
| 0.10 | 0.00 | −0.10 |
| 0.55 | 0.50 | −0.05 |
| 0.95 | 1.00 | +0.05 |
Small weight error is not the whole objective
For a linear layer, the output error is $(W-\widehat W)x$. Its size depends on the activations x, not only the weight error. Calibration methods use representative activations to choose or compensate for quantization. GPTQ and AWQ make different choices about reconstruction and salient channels; the existing quantization notes explain their base algorithms.
In a masked diffusion model, x changes with mask fraction, revealed context, sequence length, and sampler decisions. Calibrating only on clean text can miss states encountered during generation. A calibration set should cover the actual prompt/response structure and a range of denoising states, including states produced by the intended sampler when feasible.
Design the calibration experiment
Freeze the checkpoint, tokenizer, prompt template, output length, and sampler. Collect activation examples across masking levels and representative tasks. Keep calibration data separate from final evaluation. Compare clean-only calibration, randomly masked states, and trajectory-derived states under the same calibration budget. This tests coverage rather than assuming one source is best.
If calibration uses full-precision trajectories, quantized sampling can later visit different states. Measure that shift. A second calibration round using quantized trajectories is an experimental option, but it must not use held-out test answers to improve the model. Record masking and position-selection policies with the calibration artifact.
Separate local error from decision changes
| Level | Measurement | Question |
|---|---|---|
| Layer | Relative output error on held-out activations | Which operations are numerically sensitive? |
| Logits | Distribution divergence or top-token agreement | Are predictions changing at fixed states? |
| Selection | Agreement on positions chosen for reveal/remasking | Did confidence ordering change? |
| Trajectory | Validity, task accuracy, diversity, termination | What happens after decisions feed back? |
| System | Live memory, peak memory, latency, throughput | Does the deployed implementation deliver a useful operating point? |
A tiny logit perturbation can swap two nearly tied confidence values and change the reveal order. That creates a discrete difference even when average reconstruction error is small. Conversely, a visible numerical error may not change any important decision. Inspect sensitivity near decision boundaries instead of assuming every layer error has equal task impact.
Compressed storage, live memory, and speed are separate results
Weight-only quantization may dequantize into higher precision inside a kernel. Activations, temporary buffers, logits, model copies, and allocator behavior can dominate live memory. Report artifact size, resident memory after loading, and peak memory during generation separately. Measure after warm-up and synchronize asynchronous device work.
Low-bit storage does not guarantee accelerated arithmetic on the target hardware. Kernel support, fusion, group size, batch shape, and dequantization overhead affect latency. Compare configurations at matched task quality or show the full quality–latency–memory frontier. Record actual denoiser calls and tokens processed, especially if a changed sampler uses more iterations to recover quality.
What transfers to another model?
A quantization recipe, a calibration dataset, and already quantized parameters are three different objects. Test recipe transfer across checkpoints, architectures, and tasks while keeping the comparison controlled. Transfer of one successful configuration does not imply that every sensitive layer, group size, or calibration distribution remains appropriate.
Publish an experiment manifest with precision per module, calibration source and mask coverage, sampler configuration, hardware, kernels, task metrics, uncertainty, and failure examples. This makes “what compresses” a reproducible question and “what breaks” a diagnosable one. The next chapter examines another trajectory-changing intervention: starting diffusion from an autoregressive draft.
Check your understanding: Why can a low average layer error still change the final answer?
The error can change a token or position-selection decision near a tie. That decision changes later context, so the remaining denoising trajectory can diverge.
Sources and further reading
- Frantar et al.: GPTQ — Activation-informed post-training weight quantization.
- Lin et al.: AWQ — Activation-aware weight quantization.
- Sahoo et al.: simple and effective masked diffusion language models — The masked denoising setting that motivates state-aware evaluation.