← All Posts
Deep Learning · Transformers· Efficient execution

Speculative Decoding: Draft, Verify, and Correct

A cheap model proposes; the target model decides which proposals can be kept. With the correct acceptance and correction distributions, speculative sampling preserves the target distribution while sometimes producing multiple tokens per target-model invocation.

Name the two distributions at one prefix

Let $p(v)$ be the expensive target model's next-token distribution and $q(v)$ the draft model's distribution at the same accepted prefix. Draw a proposal $y\sim q$. Accept it with probability

$$a(y)=\min\!\left(1,\frac{p(y)}{q(y)}\right).$$

A sampled proposal has $q(y)>0$, so this ratio is defined on its support. If the proposal is rejected, sample a replacement from

$$r(v)=\frac{[p(v)-q(v)]_+}{\sum_u[p(u)-q(u)]_+},\qquad [z]_+=\max(z,0).$$

The residual denominator is positive whenever rejection has positive probability. If $p=q$ exactly, acceptance is certain and the residual branch is never needed. A numerical implementation still needs a sensible treatment of roundoff near that case.

Why the correction recovers the target

The probability mass of accepted proposals at value $v$ is $q(v)a(v)=\min(p(v),q(v))$. The total rejection mass is $1-\sum_v\min(p(v),q(v))$, which equals $\sum_v[p(v)-q(v)]_+$ because both distributions sum to one.

Multiplying that rejection probability by $r(v)$ contributes $[p(v)-q(v)]_+$. Adding the two paths yields

$$\min(p(v),q(v))+[p(v)-q(v)]_+=p(v).$$

This is the exact one-step distribution argument. Sampling from $p$ directly after rejection instead of from the residual generally breaks it. The draft may be inaccurate and still preserve correctness; inaccuracy primarily reduces how much work can be saved.

Three possible next tokens

Take $p=(0.6,0.3,0.1)$ and $q=(0.2,0.5,0.3)$. The acceptance probabilities are $(1,0.6,1/3)$. Accepted mass is $(0.2,0.3,0.1)$, totaling 0.6. The residual is concentrated entirely on token 1, because $[p-q]_+=(0.4,0,0)$.

The rejection path contributes another 0.4 to token 1, recovering $(0.6,0.3,0.1)$. This example makes the correction's role visible: it supplies mass the draft systematically underproduces.

Try the mechanism Watch accepted and corrected probability mass add up

After rejection, sample the normalized positive part of p − q. Sampling directly from p at that point would generally be wrong. Enable JavaScript to change the example inputs; the complete calculation remains in the article.

Verify several proposed tokens in one target pass

The draft generates a sequence of proposals autoregressively, retaining the draft distribution at each proposal prefix. The target evaluates those proposed prefixes in one causal forward pass. Walk through proposals in order, applying the acceptance test at each still-accepted prefix.

At the first rejection, sample the residual correction and discard later proposals: their prefixes include a token that is no longer part of the accepted sequence. If all proposals are accepted, the target can supply an additional token from its distribution after the final proposal, subject to stopping rules.

The verified tokens are not independent target samples. Their conditional distributions are evaluated at the appropriate prefixes, and the ordered acceptance procedure preserves that dependency structure.

Speculation needs cache rollback

The target pass may compute K/V entries for proposals that are later rejected. Those entries cannot remain as committed history. Keep the cache for the accepted prefix, discard invalid suffix entries, and process the correction under the correct context. The draft's state needs matching rollback or recomputation.

EOS and length limits can end generation before the nominal draft block is exhausted. Cache lengths, returned tokens, and stopping conditions must agree. An algorithm that returns correct text but leaves invalid cached state can fail on the following call.

When is drafting worth it?

Under a simplified constant conditional acceptance rate $a$ and a draft block of $k$ tokens, the expected number of committed tokens including correction or bonus is $1+a+\cdots+a^k$. This is an illustrative model; real acceptance varies with the prefix and position.

Compare that expected progress with the cost of drafting $k$ tokens, one target verification pass, and coordination. A larger block can waste work after early rejection. A stronger draft can improve acceptance but cost more. Throughput and latency depend on batch size and hardware as well as the probability match.

Preserve the distribution you actually intend

If temperature or top-$p$ filtering defines the desired target sampling distribution, use that resulting $p$ consistently in acceptance and correction. The recorded $q$ must likewise be the distribution that actually produced the proposal. Mixing raw logits from one policy with probabilities from another invalidates the proof.

Greedy speculative decoding has a simpler token-agreement verification rule for preserving greedy outputs, but that is not the stochastic acceptance algorithm above. Specify which behavior is being preserved.

Try it: If the draft assigns zero probability to a token that the target likes, can the output still follow the target distribution?

Yes. That token cannot arrive through an accepted proposal, but it can receive mass through the residual correction. The accepted-plus-residual decomposition recovers the target distribution.

Read and test the exact algorithm

The primary reference is Fast Inference from Transformers via Speculative Decoding. The lab verifies the one-step probability decomposition exactly for small distributions.