This article is published in English.
Speculative Decoding and Early Exit: Speeding Autoregressive Decode
Draft-then-verify and confidence-based layer exits cut decode latency when acceptance and accuracy budgets allow.
Why decode optimization sits on the critical path
LLM inference has a prefill phase that parallelizes over the prompt and a decode phase that emits tokens sequentially. Prefill can saturate GPUs; decode often cannot, because each new token depends on the previous one. That sequential dependency is why interactive latency stays expensive even when FLOPs look plentiful on paper.
Prefill phase:
Input: [token_1, token_2, ..., token_512] → all 512 tokens processed in parallel
Matrix shape: [batch, 512, 4096]
GPU utilization: high — large matrix, full Tensor Core throughput
Decode phase (one step):
Input: [token_513] → one token processed
Matrix shape: [batch, 1, 4096]
GPU utilization: low - tiny matrix, most CUDA cores idle
Attention structure during decode
At each decode step, the model forms queries against a growing key/value cache. Work per step climbs with context length, while batching opportunities are constrained by latency SLOs for chat.
One attention head, decode step:
Query Q: [8, 1, 128] → 8 × 128 = 1,024 elements
Keys K: [8, 2048, 128] → 8 × 2048 × 128 = 2M elements
Values V: [8, 2048, 128] → same
Matrix multiply (QKᵀ) per head:
Shape: [8, 1, 128] × [8, 128, 2048] → [8, 1, 2048]
FLOPs: 8 × 1 × 128 × 2048 ≈ 2.1M FLOPs per head
For 32 heads: ≈ 67M FLOPs per layer
For 32 layers: ≈ 2.1B FLOPs total per decode step
A100 peak at BF16: ~312 TFLOPS = 312 × 10¹² FLOPs/sec
Time to execute 2.1B FLOPs at 100% utilization:
2.1 × 10⁹ / 312 × 10¹² ≈ 0.0067 ms
Actual observed decode latency per step: ~10–30ms
Effective compute utilization: < 0.1%
Speculative decoding: draft then verify
A smaller draft model proposes several future tokens; the target model verifies them in one parallel forward pass, accepting a prefix and resampling on the first rejection. Accepted drafts multiply effective tokens per expensive target step.
Without speculative decoding:
Generate 5 tokens: 5 × 30ms = 150ms
With speculative decoding (K=4):
Draft 4 tokens: 4 × 1.5ms = 6ms
1 target verification pass: ~35ms (slightly longer than decode,
processes K+1=5 positions)
Expected accepted tokens per round:
(1 - 0.80^5) / (1 - 0.80) ≈ 3.36 tokens
Time per round: 6ms + 35ms = 41ms
Time per token: 41ms / 3.36 ≈ 12.2ms
Speedup: 30ms → 12.2ms ≈ 2.5×
Speedup depends on agreement between draft and target. Easy, low-entropy text accepts long drafts; surprising tokens truncate acceptance.
import torch
import torch.nn.functional as F
def speculative_decode(target_model, draft_model, input_ids,
max_new_tokens, K=4, temperature=1.0):
"""
Conceptual speculative decoding loop.
Real implementations handle KV cache management across both models.
"""
generated = input_ids.clone()
while generated.shape[1] - input_ids.shape[1] < max_new_tokens:
# --- Draft phase ---
draft_tokens = []
draft_probs = []
draft_input = generated.clone()
for _ in range(K):
with torch.no_grad():
draft_logits = draft_model(draft_input).logits[:, -1, :]
q = F.softmax(draft_logits / temperature, dim=-1)
token = torch.multinomial(q, num_samples=1)
draft_tokens.append(token)
draft_probs.append(q)
draft_input = torch.cat([draft_input, token], dim=1)
# --- Verify phase: one target forward pass over all K+1 positions ---
verify_input = torch.cat([generated] + draft_tokens, dim=1)
with torch.no_grad():
target_logits = target_model(verify_input).logits
# target_logits[:, -K-1:, :] covers all K draft positions + bonus
# --- Accept/reject ---
accepted = 0
for i in range(K):
p = F.softmax(target_logits[:, -(K+1)+i, :] / temperature, dim=-1)
q = draft_probs[i]
token = draft_tokens[i]
# Acceptance probability
accept_prob = torch.min(
torch.ones_like(p.gather(1, token)),
p.gather(1, token) / (q.gather(1, token) + 1e-9)
)
if torch.rand(1) < accept_prob:
generated = torch.cat([generated, token], dim=1)
accepted += 1
else:
# Sample corrected token and stop this round
corrected_dist = F.relu(p - q)
corrected_dist = corrected_dist / corrected_dist.sum(dim=-1, keepdim=True)
corrected_token = torch.multinomial(corrected_dist, num_samples=1)
generated = torch.cat([generated, corrected_token], dim=1)
break
else:
# All K accepted - take bonus token
bonus_logits = target_logits[:, -1, :]
p_bonus = F.softmax(bonus_logits / temperature, dim=-1)
bonus_token = torch.multinomial(p_bonus, num_samples=1)
generated = torch.cat([generated, bonus_token], dim=1)
return generated
Choose drafts from the same family when possible, distilled or quantized smaller siblings, and measure acceptance rate on your traffic—not on public blog examples.
When drafts diverge
If the draft’s distribution drifts, acceptance collapses and you pay draft cost for little gain. Monitor acceptance continuously; fall back to plain decode when it drops.
Early exit: stop deep layers when confident
Some architectures allow exiting at intermediate layers when confidence is high, saving compute on “easy” tokens.
Layer distribution of exits:Exit at layers 1–8 (very easy tokens like punctuation, articles): 15%
Exit at layers 9–16 (medium tokens, common continuations): 35%
Exit at layers 17–24 (harder tokens, named entities, numbers): 30%
Exit at layers 25–32 (full computation required): 20%Weighted average layers executed:
0.15 × 6 + 0.35 × 12 + 0.30 × 20 + 0.20 × 32
= 0.90 + 4.20 + 6.00 + 6.40
= 17.5 layers averageSpeedup vs always running 32 layers:
32 / 17.5 ≈ 1.83×
Measure accuracy impact carefully: early exit trades quality for speed and is sensitive to calibration.
Speculative vs early exit
Speculative decoding changes the token proposal loop with a second model; early exit changes depth within one model. They address related bottlenecks differently and can sometimes stack with quantization, continuous batching, and KV-cache paging.
Combined stack example:
Target: 7B model, BF16, FlashAttention, PagedAttention
→ Model: ~14 GB, memory-efficient attention, paged KV cache
Draft: 70M model, INT4, FlashAttention
→ Model: ~35 MB, near-zero memory overhead
Speculative decode:
→ with K=4, α=0.80
→ ~2.5× token generation speedup on long outputs
Full stack speedup vs FP32 no-optimization baseline:
- Quantization: 2–2.3× throughput
- FlashAttention: 20–40% attention latency reduction
- Speculative decoding: 2–2.5× decode speedup (on eligible requests)
Combined: 5–8× improvement in end-to-end tokens/second
When each helps
Speculative helps when draft agreement is high and target steps dominate latency. It fails when drafts rarely match or draft overhead exceeds gains. Early exit helps when many tokens are easy and accuracy budgets allow; it fails on hard tokens or poorly calibrated confidence.
Production perspective
Ship with metrics: acceptance rate, mean accepted length, quality eval deltas, GPU utilization, and p95 latency. Feature-flag optimizations. Keep a kill switch to vanilla decode. Remember the remaining bottleneck may be memory bandwidth, network, or client rendering—not only FLOPs—so profile before stacking every paper trick.
Closing
Prefill and decode stress hardware differently. Speculative decoding and early exit are practical levers when measured against your acceptance and accuracy curves. Treat them as production features with dashboards, not one-off benchmarks, and they will earn their keep on real traffic.
Reference scenario for intuition
Imagine a 7B target serving chat with 50–150 token replies. Prefill of a 2k-token system prompt is heavy but infrequent per turn; decode steps dominate user-perceived delay. Cutting decode step count via speculative accepts, or cutting work per step via early exit, moves the needle users feel.
Operational checklist
- Baseline vanilla p95 and quality.
- Add draft model colocated to avoid network RTT.
- Log acceptance histograms.
- A/B quality on factual suites.
- Watch CPU/GPU balance when drafts run on different devices.
- Revisit after every tokenizer or finetune change.
Common pitfalls
Using a randomly unrelated draft; ignoring temperature effects on acceptance; declaring victory from throughput benches without quality gates; stacking early exit with aggressive quantization until hallucinations spike. Each pitfall is measurable—instrument first.
Extended guidance for platform teams
Centralize inference optimization in the serving layer so app teams do not each invent draft selection. Expose headers or trace attributes showing whether speculative or early exit applied. Chargeback token bills with optimization tags so finance sees the win. Rehearse rollback. Document that speculative decoding does not remove the need for good retrieval or good prompts—it only makes generating the next tokens cheaper when the model already knows what it wants to say.
More depth on stacking
Combine continuous batching with speculative carefully: draft lengths interact with scheduler assumptions. Combine with KV cache eviction policies so long chats do not thrash. Combine with prompt caching on the prefill side so you do not “fix decode” while leaving prefill wasteful. Holistic serving design beats a single paper transplanted into a hot path.
Practitioner FAQ
Does speculative change answers? It should match the target distribution when implemented correctly; verify with paired tests. Does early exit change answers? Yes by design when it skips layers—budget the delta. Can we draft on CPU? Sometimes; measure. Is this relevant for tiny models on device? Often less than for large targets. Priority-order: fix batching and caching first, then speculative, then early exit if the stack supports it.
Reference scenario for intuition
Imagine a 7B target serving chat with 50–150 token replies. Prefill of a 2k-token system prompt is heavy but infrequent per turn; decode steps dominate user-perceived delay. Cutting decode step count via speculative accepts, or cutting work per step via early exit, moves the needle users feel.
Operational checklist
- Baseline vanilla p95 and quality.
- Add draft model colocated to avoid network RTT.
- Log acceptance histograms.
- A/B quality on factual suites.
- Watch CPU/GPU balance when drafts run on different devices.
- Revisit after every tokenizer or finetune change.
Common pitfalls
Using a randomly unrelated draft; ignoring temperature effects on acceptance; declaring victory from throughput benches without quality gates; stacking early exit with aggressive quantization until hallucinations spike. Each pitfall is measurable—instrument first.
Extended guidance for platform teams
Centralize inference optimization in the serving layer so app teams do not each invent draft selection. Expose headers or trace attributes showing whether speculative or early exit applied. Chargeback token bills with optimization tags so finance sees the win. Rehearse rollback. Document that speculative decoding does not remove the need for good retrieval or good prompts—it only makes generating the next tokens cheaper when the model already knows what it wants to say.
More depth on stacking
Combine continuous batching with speculative carefully: draft lengths interact with scheduler assumptions. Combine with KV cache eviction policies so long chats do not thrash. Combine with prompt caching on the prefill side so you do not “fix decode” while leaving prefill wasteful. Holistic serving design beats a single paper transplanted into a hot path.
Practitioner FAQ
Does speculative change answers? It should match the target distribution when implemented correctly; verify with paired tests. Does early exit change answers? Yes by design when it skips layers—budget the delta. Can we draft on CPU? Sometimes; measure. Is this relevant for tiny models on device? Often less than for large targets. Priority-order: fix batching and caching first, then speculative, then early exit if the stack supports it.
Worked numeric intuition
Suppose a target step costs 10 ms and a draft proposes 5 tokens with 60% average acceptance of 3 tokens. Effective cost per accepted token falls versus vanilla one-token steps, even after draft overhead, when acceptance stays healthy. If acceptance falls to ~1 token, the scheme loses. That sensitivity is why dashboards beat anecdotes.
Quality regression protocol
Before enabling globally, run fixed prompts across factual QA, coding, and refusal suites. Compare token-identical rates when speculative is configured for exact distribution match. Investigate any systematic drift. For early exit, track win-rate on graded tasks and human preference where available.
Hardware placement
Colocate draft and target on the same node when possible. Cross-host drafts add network jitter that can erase gains. Watch memory: two models plus KV cache can OOM a box that comfortably held one.
Scheduling interactions
Continuous batching servers must account for variable speculative expands. Poor schedulers fragment batches and hurt utilization. Coordinate with serving maintainers; do not flip flags only in app code.
Remaining bottleneck honesty
After decode improves, users may still wait on tool calls, retrieval, or client-side markdown. Trace end-to-end. Optimization theater on the wrong span wastes engineering time.
Summary reprise
Sequential decode is the structural tax of autoregression. Speculative decoding and early exit reduce that tax under measurable conditions. Ship them with the same discipline as any production feature: metrics, flags, rollbacks, and clear owners in the serving platform team.
Worked numeric intuition
Suppose a target step costs 10 ms and a draft proposes 5 tokens with 60% average acceptance of 3 tokens. Effective cost per accepted token falls versus vanilla one-token steps, even after draft overhead, when acceptance stays healthy. If acceptance falls to ~1 token, the scheme loses. That sensitivity is why dashboards beat anecdotes.
Quality regression protocol
Before enabling globally, run fixed prompts across factual QA, coding, and refusal suites. Compare token-identical rates when speculative is configured for exact distribution match. Investigate any systematic drift. For early exit, track win-rate on graded tasks and human preference where available.
Hardware placement
Colocate draft and target on the same node when possible. Cross-host drafts add network jitter that can erase gains. Watch memory: two models plus KV cache can OOM a box that comfortably held one.
Scheduling interactions
Continuous batching servers must account for variable speculative expands. Poor schedulers fragment batches and hurt utilization. Coordinate with serving maintainers; do not flip flags only in app code.
Remaining bottleneck honesty
After decode improves, users may still wait on tool calls, retrieval, or client-side markdown. Trace end-to-end. Optimization theater on the wrong span wastes engineering time.
Summary reprise
Sequential decode is the structural tax of autoregression. Speculative decoding and early exit reduce that tax under measurable conditions. Ship them with the same discipline as any production feature: metrics, flags, rollbacks, and clear owners in the serving platform team.
Worked numeric intuition
Suppose a target step costs 10 ms and a draft proposes 5 tokens with 60% average acceptance of 3 tokens. Effective cost per accepted token falls versus vanilla one-token steps, even after draft overhead, when acceptance stays healthy. If acceptance falls to ~1 token, the scheme loses. That sensitivity is why dashboards beat anecdotes.
Quality regression protocol
Before enabling globally, run fixed prompts across factual QA, coding, and refusal suites. Compare token-identical rates when speculative is configured for exact distribution match. Investigate any systematic drift. For early exit, track win-rate on graded tasks and human preference where available.
Hardware placement
Colocate draft and target on the same node when possible. Cross-host drafts add network jitter that can erase gains. Watch memory: two models plus KV cache can OOM a box that comfortably held one.
Scheduling interactions
Continuous batching servers must account for variable speculative expands. Poor schedulers fragment batches and hurt utilization. Coordinate with serving maintainers; do not flip flags only in app code.
Remaining bottleneck honesty
After decode improves, users may still wait on tool calls, retrieval, or client-side markdown. Trace end-to-end. Optimization theater on the wrong span wastes engineering time.
Summary reprise
Sequential decode is the structural tax of autoregression. Speculative decoding and early exit reduce that tax under measurable conditions. Ship them with the same discipline as any production feature: metrics, flags, rollbacks, and clear owners in the serving platform team.
Worked numeric intuition
Suppose a target step costs 10 ms and a draft proposes 5 tokens with 60% average acceptance of 3 tokens. Effective cost per accepted token falls versus vanilla one-token steps, even after draft overhead, when acceptance stays healthy. If acceptance falls to ~1 token, the scheme loses. That sensitivity is why dashboards beat anecdotes.
Quality regression protocol
Before enabling globally, run fixed prompts across factual QA, coding, and refusal suites. Compare token-identical rates when speculative is configured for exact distribution match. Investigate any systematic drift. For early exit, track win-rate on graded tasks and human preference where available.
Hardware placement
Colocate draft and target on the same node when possible. Cross-host drafts add network jitter that can erase gains. Watch memory: two models plus KV cache can OOM a box that comfortably held one.
Scheduling interactions
Continuous batching servers must account for variable speculative expands. Poor schedulers fragment batches and hurt utilization. Coordinate with serving maintainers; do not flip flags only in app code.
Remaining bottleneck honesty
After decode improves, users may still wait on tool calls, retrieval, or client-side markdown. Trace end-to-end. Optimization theater on the wrong span wastes engineering time.
Summary reprise
Sequential decode is the structural tax of autoregression. Speculative decoding and early exit reduce that tax under measurable conditions. Ship them with the same discipline as any production feature: metrics, flags, rollbacks, and clear owners in the serving platform team.