← All Posts
Deep Learning · Diffusion & Flow Models · Large Language Diffusion Models

Speculative Refinement: Give Diffusion a Draft, Then Evaluate the Edits

Before you start: Sampling and remasking and diffusion evaluation.

An all-mask canvas asks a diffusion model to discover both the structure and the details of an answer. An autoregressive draft supplies a possible structure first. Speculative Refinement, or SpecRef, selectively masks uncertain parts of that draft and lets a masked diffusion model fill them again. The interesting evaluation question is whether the edits repair mistakes without damaging what was already correct.

01AR draftGenerate text and record token distributions
02Align and selectMap uncertainty to refiner tokens
03MaskRemove selected spans while preserving context
04RefineRun the masked diffusion sampler
05Evaluate editsCount repairs and newly introduced errors
This is a generation strategy that changes the output distribution. It is not the rejection-corrected verifier used in exact autoregressive speculative decoding.

Use uncertainty to choose what to revisit

$$H_i=-\sum_vq_\phi(v\mid c,x_{

The entropy uses the entire next-token distribution, not just the probability of the sampled token. A high value indicates uncertainty across alternatives. SpecRef ranks positions by entropy, maps them to the refiner’s tokenization, masks a selected fraction, and adds task-specific mask expansion. High entropy is a heuristic for where editing may help; a confidently wrong token can remain untouched.

The refiner sees “return x [MASK] [MASK]”. Whether an edit is correct depends on the function’s specification, not on confidence alone.
Draft tokenIllustrative entropyAction at this threshold
return0.1Keep
x0.2Keep
+1.3Mask
10.9Mask

Tokenizers do not share position numbers

A drafter and refiner can split identical text differently. Decode the draft, obtain character offsets for both tokenizations, and map uncertainty over overlapping spans using a declared aggregation rule. Preserve prompt boundaries and special tokens. Copying entropy at drafter index 17 to refiner index 17 is generally invalid.

The published method uses top-percentile entropy masking with additional math-neighborhood expansion and a tail-masking heuristic. Those extensions can make the realized mask fraction larger than the nominal percentile. Record the actual selected positions and realized fraction in every run. They are part of the method’s behavior, not incidental preprocessing.

Selective corruption changes the starting distribution

Random forward masking of clean training text and entropy-selected masking of a possibly incorrect draft are different distributions. The denoiser may still be a useful editor, but this handoff does not inherit exact sampling claims merely by calling the intermediate state a diffusion time. Fixed draft tokens constrain what the refiner can correct.

At zero selected positions, the simplest preserve-unmasked policy returns the draft. At all response positions masked, it approaches a full-mask start for the declared canvas and sampler, while still paying any draft-generation cost. Intermediate ratios trade preserved structure against editing freedom. Compare them empirically rather than assuming more refinement always helps.

Measure correction and corruption separately

Use task-level outcomes, claim-level outcomes, or aligned token outcomes as appropriate. State the unit; token correctness is not always well-defined.
Before / afterCorrect afterWrong after
Correct beforeRetained successesIntroduced failures
Wrong beforeSuccessful repairsRemaining failures

Suppose a toy task set has 60 correct drafts out of 100. Refinement repairs 15 of the 40 failures but breaks 8 of the 60 successes. Final accuracy is (60−8+15)/100=67%. Repair rate is 15/40=37.5%; corruption rate is 8/60≈13.3%. Reporting only the repair rate would hide a major cost. This transition table makes the refinement tension concrete.

Benchmark protocols can measure different capabilities

For code, measure execution-based correctness and report syntax validity separately. A valid scaffold can improve parsing without fixing logic. Preserve raw generations and use a generator-appropriate extraction rule; blindly truncating or cleaning non-AR output can invalidate results. For math, define final-answer extraction and exact-match normalization. Likelihood-based multiple-choice scoring asks a different question from actually generating an answer.

Compare AR draft alone, diffusion alone, and the hybrid with matched prompts and documented budgets. Include random masking versus entropy masking and the expansion heuristics as ablations. Evaluate both matched-call and matched-wall-clock settings when useful, because an AR token step and a full diffusion pass have different costs.

$$T_{\mathrm{total}}=T_{\mathrm{draft}}+T_{\mathrm{alignment}}+T_{\mathrm{masking}}+T_{\mathrm{refinement}}+T_{\mathrm{postprocess}}.$$

Measure each term rather than assuming drafting is negligible. Include memory when both models coexist, and show task quality against total latency. A hybrid can outperform one component while still losing to the other at a particular budget; the comparison must make that visible.

Check your understanding: Does keeping a draft token guarantee it was correct?

No. The preservation rule reflects a confidence heuristic and the chosen mask budget. Confidently wrong draft tokens are an important failure slice.

Sources and further reading