Models · reviewed · reviewed Oct 6, 2026 · 4 min
How does speculative decoding work?
A draft proposes tokens; the target scores their conditional prefixes in a batch. A probability-aware rule accepts a prefix, corrects the first rejection, and discards later proposals. Exact speculative sampling preserves the target distribution; speed depends on draft cost and execution.
Predict ahead without committing ahead
Autoregressive generation normally produces a token, appends it to the context, and uses that new prefix to produce the next token. Those dependencies make the sequence difficult to generate through independent parallel guesses.
Speculative decoding introduces a draft mechanism that proposes a short continuation cheaply. The target model then scores several conditional positions together. It may retain several proposals from one target batch, instead of requiring a separate target pass for every retained token.
The target still owns the intended sampling distribution. The draft supplies candidates; its output is not automatically published. The verification rule here concerns probabilities and prefix consistency, not whether a sentence is factually correct.
Follow one block into committed output
The experiment has three toy token IDs, A, B, and C. Authored tables define the next-token probabilities for an empty prefix and for prefixes ending in each ID. The aligned draft exactly matches the target tables; the biased draft strongly favors A.
Draft three tokens, inspect the uncommitted proposal, then verify and commit. With the aligned draft, all proposals are accepted and the target can add a bonus token if the output limit allows it. Switch to the biased draft and seed 1: the first block exposes a rejection and discards its later proposals.
Explore the mechanism
Draft a block, then verify its prefix
Committed output
No committed tokens
Draft workspace
Draft a block to see the proposal before it can change the committed output.
- Committed tokens
- 0 / 8
- Target batch passes
- 0
- Draft token evaluations
- 0
The committed prefix is the input to the next draft block.
Inspect target and draft probabilities
Verify a draft to inspect its target and draft tables. The token order is A, B, C.
A/B/C are toy token IDs, and conditional probability tables are authored. Sampling, acceptance draws, residual correction, and state transitions are computed with a fixed seed. The aligned draft exactly matches the target table. No neural model, parallel GPU execution, latency, or speedup is measured; counters describe this algorithmic trace.
After a rejection, the prefix that would have preceded later draft tokens is no longer valid. Those later proposals must be discarded, even if one looks plausible in isolation. The next draft begins from the corrected, committed output.
Changing the draft, width, or seed resets the trace. This toy stops at eight committed tokens and has no end-of-sequence token. Its work counters are algorithmic counts, not measured latency or LLM performance.
Acceptance needs the probability ratio
At one position, call the target distribution p and the draft distribution q. A token sampled from q is accepted with probability min(1, p(token)/q(token)). A uniform draw makes that decision.
If it is rejected, the replacement comes from the normalized positive part of p − q. Sampling directly from p at that point would ignore the mass already contributed by accepted draft samples and generally bias the result.
The inspector exposes probabilities, draws, and the correction distribution. With an empty prefix, the biased draft assigns [0.80, 0.15, 0.05], while the target assigns [0.50, 0.30, 0.20]. The residual correction distribution is [0, 0.50, 0.50]; it cannot select A in that rejected-position case.
When every draft token is accepted, an additional token can be sampled from the target after that accepted prefix. The output limit can remove the need for this bonus at the final boundary.
The same distribution need not mean the same trace
Exact speculative sampling preserves the target’s probability distribution under its assumptions. It does not imply that every implementation using the same numeric seed produces the same individual text as an ordinary decoder. Different procedures can consume random draws differently; numerical execution can also differ.
The earlier blockwise parallel decoding work studies a greedy variant: verify agreement with the greedy model and retain the matching prefix. That mechanism should not be confused with the probability-correct rejection rule for stochastic sampling.
Likewise, “the target accepted this token” does not mean it checked a citation, proved an argument, or approved a tool effect. Speculation accelerates a decoding procedure. It does not add an independent truth oracle.
Where the speedup can disappear
A longer draft creates more opportunities to retain tokens per target batch. It also creates more draft work and can leave more discarded computation after an early rejection. Acceptance behaviour, draft cost, hardware, memory access, batching, and runtime implementation determine the operating point.
The widget groups target-table lookups into a symbolic batch pass; it does not run a transformer on a GPU. Four committed tokens after one displayed pass demonstrate an algorithmic possibility, not a four-times speedup.
A distilled model can serve as a draft, but distillation and speculative decoding answer different questions. Distillation trains another model. Speculation uses proposals during inference while keeping the target distribution as the reference.
Why the residual restores the target distribution
For one position, the probability mass of an accepted draft token is q(token) × min(1, p(token)/q(token)) = min(p(token), q(token)).
The remaining mass is p(token) − min(p(token), q(token)). Normalizing the positive residual p − q and sampling it after rejection supplies exactly that remainder. Accepted and corrected mass therefore sum to p(token).
The toy target uses [0.50, 0.30, 0.20] initially; after A, B, or C it uses [0.20, 0.60, 0.20], [0.55, 0.15, 0.30], or [0.25, 0.25, 0.50]. These are authored conditional probabilities, not measured model outputs. Seeds drive a deterministic local pseudorandom generator; no training or remote inference occurs.
Sources
Sources and further reading
- 01Fast Inference from Transformers via Speculative DecodingLeviathan, Kalman, and Matias · research · published Nov 30, 2022 · source checked Oct 6, 2026
Primary exact speculative-sampling algorithm: draft proposals, target batch scoring, probability-ratio acceptance, residual correction, and conditional speed trade-offs.
- 02Accelerating Large Language Model Decoding with Speculative SamplingChen et al. · research · published Feb 2, 2023 · source checked Oct 6, 2026
Independent draft/target speculative sampling work explains distribution preservation and the importance of execution and draft overhead.
- 03Blockwise Parallel Decoding for Deep Autoregressive ModelsStern, Shazeer, and Uszkoreit · research · published Nov 7, 2018 · source checked Oct 6, 2026
Greedy block prediction and longest-prefix verification provide a precursor distinct from probability-correct stochastic speculation.
