Speculative Refinement: Give Diffusion a Draft, Then Evaluate the Edits
Before you start: Sampling and remasking and diffusion evaluation.
An all-mask canvas asks a diffusion model to discover both the structure and the details of an answer. An autoregressive draft supplies a possible structure first. Speculative Refinement, or SpecRef, selectively masks uncertain parts of that draft and lets a masked diffusion model fill them again. The interesting evaluation question is whether the edits repair mistakes without damaging what was already correct.
Use uncertainty to choose what to revisit
The entropy uses the entire next-token distribution, not just the probability of the sampled token. A high value indicates uncertainty across alternatives. SpecRef ranks positions by entropy, maps them to the refiner’s tokenization, masks a selected fraction, and adds task-specific mask expansion. High entropy is a heuristic for where editing may help; a confidently wrong token can remain untouched.
| Draft token | Illustrative entropy | Action at this threshold |
|---|---|---|
| return | 0.1 | Keep |
| x | 0.2 | Keep |
| + | 1.3 | Mask |
| 1 | 0.9 | Mask |
Tokenizers do not share position numbers
A drafter and refiner can split identical text differently. Decode the draft, obtain character offsets for both tokenizations, and map uncertainty over overlapping spans using a declared aggregation rule. Preserve prompt boundaries and special tokens. Copying entropy at drafter index 17 to refiner index 17 is generally invalid.
The published method uses top-percentile entropy masking with additional math-neighborhood expansion and a tail-masking heuristic. Those extensions can make the realized mask fraction larger than the nominal percentile. Record the actual selected positions and realized fraction in every run. They are part of the method’s behavior, not incidental preprocessing.
Selective corruption changes the starting distribution
Random forward masking of clean training text and entropy-selected masking of a possibly incorrect draft are different distributions. The denoiser may still be a useful editor, but this handoff does not inherit exact sampling claims merely by calling the intermediate state a diffusion time. Fixed draft tokens constrain what the refiner can correct.
At zero selected positions, the simplest preserve-unmasked policy returns the draft. At all response positions masked, it approaches a full-mask start for the declared canvas and sampler, while still paying any draft-generation cost. Intermediate ratios trade preserved structure against editing freedom. Compare them empirically rather than assuming more refinement always helps.
Measure correction and corruption separately
| Before / after | Correct after | Wrong after |
|---|---|---|
| Correct before | Retained successes | Introduced failures |
| Wrong before | Successful repairs | Remaining failures |
Suppose a toy task set has 60 correct drafts out of 100. Refinement repairs 15 of the 40 failures but breaks 8 of the 60 successes. Final accuracy is (60−8+15)/100=67%. Repair rate is 15/40=37.5%; corruption rate is 8/60≈13.3%. Reporting only the repair rate would hide a major cost. This transition table makes the refinement tension concrete.
Benchmark protocols can measure different capabilities
For code, measure execution-based correctness and report syntax validity separately. A valid scaffold can improve parsing without fixing logic. Preserve raw generations and use a generator-appropriate extraction rule; blindly truncating or cleaning non-AR output can invalidate results. For math, define final-answer extraction and exact-match normalization. Likelihood-based multiple-choice scoring asks a different question from actually generating an answer.
Compare AR draft alone, diffusion alone, and the hybrid with matched prompts and documented budgets. Include random masking versus entropy masking and the expansion heuristics as ablations. Evaluate both matched-call and matched-wall-clock settings when useful, because an AR token step and a full diffusion pass have different costs.
Measure each term rather than assuming drafting is negligible. Include memory when both models coexist, and show task quality against total latency. A hybrid can outperform one component while still losing to the other at a particular budget; the comparison must make that visible.
Check your understanding: Does keeping a draft token guarantee it was correct?
No. The preservation rule reflects a confidence heuristic and the chosen mask budget. Confidently wrong draft tokens are an important failure slice.
Sources and further reading
- Gupta, Mishra, Trivedi & Kumar: Speculative Refinement — The hybrid method and its evaluation across generation protocols.
- SpecRef full paper — Entropy masking, alignment, expansion heuristics, and experimental setup.
- Leviathan et al.: fast inference from transformers via speculative decoding — The separate rejection-corrected autoregressive sampling method.