Home / Articles / Speculative Decoding and Early Exit: Speeding Autoregressive Decode

This article is published in English.

Speculative Decoding and Early Exit: Speeding Autoregressive Decode

Draft-then-verify and confidence-based layer exits cut decode latency when acceptance and accuracy budgets allow.

2856 words

Why decode optimization sits on the critical path

LLM inference has a prefill phase that parallelizes over the prompt and a decode phase that emits tokens sequentially. Prefill can saturate GPUs; decode often cannot, because each new token depends on the previous one. That sequential dependency is why interactive latency stays expensive even when FLOPs look plentiful on paper.

Prefill phase:
  Input: [token_1, token_2, ..., token_512]  → all 512 tokens processed in parallel
  Matrix shape: [batch, 512, 4096]
  GPU utilization: high — large matrix, full Tensor Core throughput

Decode phase (one step):
  Input: [token_513]  → one token processed
  Matrix shape: [batch, 1, 4096]
  GPU utilization: low - tiny matrix, most CUDA cores idle

Attention structure during decode

At each decode step, the model forms queries against a growing key/value cache. Work per step climbs with context length, while batching opportunities are constrained by latency SLOs for chat.

One attention head, decode step:

Query Q: [8, 1, 128]   →  8 × 128 = 1,024 elements
Keys  K: [8, 2048, 128] →  8 × 2048 × 128 = 2M elements
Values V: [8, 2048, 128] →  same

Matrix multiply (QKᵀ) per head:
  Shape: [8, 1, 128] × [8, 128, 2048] → [8, 1, 2048]
  FLOPs: 8 × 1 × 128 × 2048 ≈ 2.1M FLOPs per head

For 32 heads: ≈ 67M FLOPs per layer
For 32 layers: ≈ 2.1B FLOPs total per decode step

A100 peak at BF16: ~312 TFLOPS = 312 × 10¹² FLOPs/sec

Time to execute 2.1B FLOPs at 100% utilization:
  2.1 × 10⁹ / 312 × 10¹² ≈ 0.0067 ms

Actual observed decode latency per step: ~10–30ms

Effective compute utilization: < 0.1%

Speculative decoding: draft then verify

A smaller draft model proposes several future tokens; the target model verifies them in one parallel forward pass, accepting a prefix and resampling on the first rejection. Accepted drafts multiply effective tokens per expensive target step.

Without speculative decoding:
  Generate 5 tokens: 5 × 30ms = 150ms

With speculative decoding (K=4):
  Draft 4 tokens: 4 × 1.5ms = 6ms
  1 target verification pass: ~35ms  (slightly longer than decode,
                                       processes K+1=5 positions)
  Expected accepted tokens per round:
    (1 - 0.80^5) / (1 - 0.80) ≈ 3.36 tokens
  Time per round: 6ms + 35ms = 41ms
  Time per token: 41ms / 3.36 ≈ 12.2ms
Speedup: 30ms → 12.2ms ≈ 2.5×

Speedup depends on agreement between draft and target. Easy, low-entropy text accepts long drafts; surprising tokens truncate acceptance.

import torch
import torch.nn.functional as F

def speculative_decode(target_model, draft_model, input_ids,
                       max_new_tokens, K=4, temperature=1.0):
    """
    Conceptual speculative decoding loop.
    Real implementations handle KV cache management across both models.
    """
    generated = input_ids.clone()
    while generated.shape[1] - input_ids.shape[1] < max_new_tokens:
        # --- Draft phase ---
        draft_tokens = []
        draft_probs = []
        draft_input = generated.clone()
        for _ in range(K):
            with torch.no_grad():
                draft_logits = draft_model(draft_input).logits[:, -1, :]
            q = F.softmax(draft_logits / temperature, dim=-1)
            token = torch.multinomial(q, num_samples=1)
            draft_tokens.append(token)
            draft_probs.append(q)
            draft_input = torch.cat([draft_input, token], dim=1)
        # --- Verify phase: one target forward pass over all K+1 positions ---
        verify_input = torch.cat([generated] + draft_tokens, dim=1)
        with torch.no_grad():
            target_logits = target_model(verify_input).logits
        # target_logits[:, -K-1:, :] covers all K draft positions + bonus
        # --- Accept/reject ---
        accepted = 0
        for i in range(K):
            p = F.softmax(target_logits[:, -(K+1)+i, :] / temperature, dim=-1)
            q = draft_probs[i]
            token = draft_tokens[i]
            # Acceptance probability
            accept_prob = torch.min(
                torch.ones_like(p.gather(1, token)),
                p.gather(1, token) / (q.gather(1, token) + 1e-9)
            )
            if torch.rand(1) < accept_prob:
                generated = torch.cat([generated, token], dim=1)
                accepted += 1
            else:
                # Sample corrected token and stop this round
                corrected_dist = F.relu(p - q)
                corrected_dist = corrected_dist / corrected_dist.sum(dim=-1, keepdim=True)
                corrected_token = torch.multinomial(corrected_dist, num_samples=1)
                generated = torch.cat([generated, corrected_token], dim=1)
                break
        else:
            # All K accepted - take bonus token
            bonus_logits = target_logits[:, -1, :]
            p_bonus = F.softmax(bonus_logits / temperature, dim=-1)
            bonus_token = torch.multinomial(p_bonus, num_samples=1)
            generated = torch.cat([generated, bonus_token], dim=1)
    return generated

Choose drafts from the same family when possible, distilled or quantized smaller siblings, and measure acceptance rate on your traffic—not on public blog examples.

When drafts diverge

If the draft’s distribution drifts, acceptance collapses and you pay draft cost for little gain. Monitor acceptance continuously; fall back to plain decode when it drops.

Early exit: stop deep layers when confident

Some architectures allow exiting at intermediate layers when confidence is high, saving compute on “easy” tokens.

Layer distribution of exits:Exit at layers 1–8  (very easy tokens like punctuation, articles):  15%
Exit at layers 9–16 (medium tokens, common continuations):           35%
Exit at layers 17–24 (harder tokens, named entities, numbers):       30%
Exit at layers 25–32 (full computation required):                    20%Weighted average layers executed:
  0.15 × 6 + 0.35 × 12 + 0.30 × 20 + 0.20 × 32
  = 0.90 + 4.20 + 6.00 + 6.40
  = 17.5 layers averageSpeedup vs always running 32 layers:
  32 / 17.5 ≈ 1.83×

Measure accuracy impact carefully: early exit trades quality for speed and is sensitive to calibration.

Speculative vs early exit

Speculative decoding changes the token proposal loop with a second model; early exit changes depth within one model. They address related bottlenecks differently and can sometimes stack with quantization, continuous batching, and KV-cache paging.

Combined stack example:

Target: 7B model, BF16, FlashAttention, PagedAttention
  → Model: ~14 GB, memory-efficient attention, paged KV cache

Draft: 70M model, INT4, FlashAttention
  → Model: ~35 MB, near-zero memory overhead

Speculative decode:
  → with K=4, α=0.80
  → ~2.5× token generation speedup on long outputs

Full stack speedup vs FP32 no-optimization baseline:
  - Quantization: 2–2.3× throughput
  - FlashAttention: 20–40% attention latency reduction
  - Speculative decoding: 2–2.5× decode speedup (on eligible requests)

Combined: 5–8× improvement in end-to-end tokens/second

When each helps

Speculative helps when draft agreement is high and target steps dominate latency. It fails when drafts rarely match or draft overhead exceeds gains. Early exit helps when many tokens are easy and accuracy budgets allow; it fails on hard tokens or poorly calibrated confidence.

Production perspective

Ship with metrics: acceptance rate, mean accepted length, quality eval deltas, GPU utilization, and p95 latency. Feature-flag optimizations. Keep a kill switch to vanilla decode. Remember the remaining bottleneck may be memory bandwidth, network, or client rendering—not only FLOPs—so profile before stacking every paper trick.

Closing

Prefill and decode stress hardware differently. Speculative decoding and early exit are practical levers when measured against your acceptance and accuracy curves. Treat them as production features with dashboards, not one-off benchmarks, and they will earn their keep on real traffic.

Reference scenario for intuition

Imagine a 7B target serving chat with 50–150 token replies. Prefill of a 2k-token system prompt is heavy but infrequent per turn; decode steps dominate user-perceived delay. Cutting decode step count via speculative accepts, or cutting work per step via early exit, moves the needle users feel.

Operational checklist

  • Baseline vanilla p95 and quality.
  • Add draft model colocated to avoid network RTT.
  • Log acceptance histograms.
  • A/B quality on factual suites.
  • Watch CPU/GPU balance when drafts run on different devices.
  • Revisit after every tokenizer or finetune change.

Common pitfalls

Using a randomly unrelated draft; ignoring temperature effects on acceptance; declaring victory from throughput benches without quality gates; stacking early exit with aggressive quantization until hallucinations spike. Each pitfall is measurable—instrument first.

Extended guidance for platform teams

Centralize inference optimization in the serving layer so app teams do not each invent draft selection. Expose headers or trace attributes showing whether speculative or early exit applied. Chargeback token bills with optimization tags so finance sees the win. Rehearse rollback. Document that speculative decoding does not remove the need for good retrieval or good prompts—it only makes generating the next tokens cheaper when the model already knows what it wants to say.

More depth on stacking

Combine continuous batching with speculative carefully: draft lengths interact with scheduler assumptions. Combine with KV cache eviction policies so long chats do not thrash. Combine with prompt caching on the prefill side so you do not “fix decode” while leaving prefill wasteful. Holistic serving design beats a single paper transplanted into a hot path.

Practitioner FAQ

Does speculative change answers? It should match the target distribution when implemented correctly; verify with paired tests. Does early exit change answers? Yes by design when it skips layers—budget the delta. Can we draft on CPU? Sometimes; measure. Is this relevant for tiny models on device? Often less than for large targets. Priority-order: fix batching and caching first, then speculative, then early exit if the stack supports it.

Reference scenario for intuition

Imagine a 7B target serving chat with 50–150 token replies. Prefill of a 2k-token system prompt is heavy but infrequent per turn; decode steps dominate user-perceived delay. Cutting decode step count via speculative accepts, or cutting work per step via early exit, moves the needle users feel.

Operational checklist

  • Baseline vanilla p95 and quality.
  • Add draft model colocated to avoid network RTT.
  • Log acceptance histograms.
  • A/B quality on factual suites.
  • Watch CPU/GPU balance when drafts run on different devices.
  • Revisit after every tokenizer or finetune change.

Common pitfalls

Using a randomly unrelated draft; ignoring temperature effects on acceptance; declaring victory from throughput benches without quality gates; stacking early exit with aggressive quantization until hallucinations spike. Each pitfall is measurable—instrument first.

Extended guidance for platform teams

Centralize inference optimization in the serving layer so app teams do not each invent draft selection. Expose headers or trace attributes showing whether speculative or early exit applied. Chargeback token bills with optimization tags so finance sees the win. Rehearse rollback. Document that speculative decoding does not remove the need for good retrieval or good prompts—it only makes generating the next tokens cheaper when the model already knows what it wants to say.

More depth on stacking

Combine continuous batching with speculative carefully: draft lengths interact with scheduler assumptions. Combine with KV cache eviction policies so long chats do not thrash. Combine with prompt caching on the prefill side so you do not “fix decode” while leaving prefill wasteful. Holistic serving design beats a single paper transplanted into a hot path.

Practitioner FAQ

Does speculative change answers? It should match the target distribution when implemented correctly; verify with paired tests. Does early exit change answers? Yes by design when it skips layers—budget the delta. Can we draft on CPU? Sometimes; measure. Is this relevant for tiny models on device? Often less than for large targets. Priority-order: fix batching and caching first, then speculative, then early exit if the stack supports it.

Worked numeric intuition

Suppose a target step costs 10 ms and a draft proposes 5 tokens with 60% average acceptance of 3 tokens. Effective cost per accepted token falls versus vanilla one-token steps, even after draft overhead, when acceptance stays healthy. If acceptance falls to ~1 token, the scheme loses. That sensitivity is why dashboards beat anecdotes.

Quality regression protocol

Before enabling globally, run fixed prompts across factual QA, coding, and refusal suites. Compare token-identical rates when speculative is configured for exact distribution match. Investigate any systematic drift. For early exit, track win-rate on graded tasks and human preference where available.

Hardware placement

Colocate draft and target on the same node when possible. Cross-host drafts add network jitter that can erase gains. Watch memory: two models plus KV cache can OOM a box that comfortably held one.

Scheduling interactions

Continuous batching servers must account for variable speculative expands. Poor schedulers fragment batches and hurt utilization. Coordinate with serving maintainers; do not flip flags only in app code.

Remaining bottleneck honesty

After decode improves, users may still wait on tool calls, retrieval, or client-side markdown. Trace end-to-end. Optimization theater on the wrong span wastes engineering time.

Summary reprise

Sequential decode is the structural tax of autoregression. Speculative decoding and early exit reduce that tax under measurable conditions. Ship them with the same discipline as any production feature: metrics, flags, rollbacks, and clear owners in the serving platform team.

Worked numeric intuition

Suppose a target step costs 10 ms and a draft proposes 5 tokens with 60% average acceptance of 3 tokens. Effective cost per accepted token falls versus vanilla one-token steps, even after draft overhead, when acceptance stays healthy. If acceptance falls to ~1 token, the scheme loses. That sensitivity is why dashboards beat anecdotes.

Quality regression protocol

Before enabling globally, run fixed prompts across factual QA, coding, and refusal suites. Compare token-identical rates when speculative is configured for exact distribution match. Investigate any systematic drift. For early exit, track win-rate on graded tasks and human preference where available.

Hardware placement

Colocate draft and target on the same node when possible. Cross-host drafts add network jitter that can erase gains. Watch memory: two models plus KV cache can OOM a box that comfortably held one.

Scheduling interactions

Continuous batching servers must account for variable speculative expands. Poor schedulers fragment batches and hurt utilization. Coordinate with serving maintainers; do not flip flags only in app code.

Remaining bottleneck honesty

After decode improves, users may still wait on tool calls, retrieval, or client-side markdown. Trace end-to-end. Optimization theater on the wrong span wastes engineering time.

Summary reprise

Sequential decode is the structural tax of autoregression. Speculative decoding and early exit reduce that tax under measurable conditions. Ship them with the same discipline as any production feature: metrics, flags, rollbacks, and clear owners in the serving platform team.

Worked numeric intuition

Suppose a target step costs 10 ms and a draft proposes 5 tokens with 60% average acceptance of 3 tokens. Effective cost per accepted token falls versus vanilla one-token steps, even after draft overhead, when acceptance stays healthy. If acceptance falls to ~1 token, the scheme loses. That sensitivity is why dashboards beat anecdotes.

Quality regression protocol

Before enabling globally, run fixed prompts across factual QA, coding, and refusal suites. Compare token-identical rates when speculative is configured for exact distribution match. Investigate any systematic drift. For early exit, track win-rate on graded tasks and human preference where available.

Hardware placement

Colocate draft and target on the same node when possible. Cross-host drafts add network jitter that can erase gains. Watch memory: two models plus KV cache can OOM a box that comfortably held one.

Scheduling interactions

Continuous batching servers must account for variable speculative expands. Poor schedulers fragment batches and hurt utilization. Coordinate with serving maintainers; do not flip flags only in app code.

Remaining bottleneck honesty

After decode improves, users may still wait on tool calls, retrieval, or client-side markdown. Trace end-to-end. Optimization theater on the wrong span wastes engineering time.

Summary reprise

Sequential decode is the structural tax of autoregression. Speculative decoding and early exit reduce that tax under measurable conditions. Ship them with the same discipline as any production feature: metrics, flags, rollbacks, and clear owners in the serving platform team.

Worked numeric intuition

Suppose a target step costs 10 ms and a draft proposes 5 tokens with 60% average acceptance of 3 tokens. Effective cost per accepted token falls versus vanilla one-token steps, even after draft overhead, when acceptance stays healthy. If acceptance falls to ~1 token, the scheme loses. That sensitivity is why dashboards beat anecdotes.

Quality regression protocol

Before enabling globally, run fixed prompts across factual QA, coding, and refusal suites. Compare token-identical rates when speculative is configured for exact distribution match. Investigate any systematic drift. For early exit, track win-rate on graded tasks and human preference where available.

Hardware placement

Colocate draft and target on the same node when possible. Cross-host drafts add network jitter that can erase gains. Watch memory: two models plus KV cache can OOM a box that comfortably held one.

Scheduling interactions

Continuous batching servers must account for variable speculative expands. Poor schedulers fragment batches and hurt utilization. Coordinate with serving maintainers; do not flip flags only in app code.

Remaining bottleneck honesty

After decode improves, users may still wait on tool calls, retrieval, or client-side markdown. Trace end-to-end. Optimization theater on the wrong span wastes engineering time.

Summary reprise

Sequential decode is the structural tax of autoregression. Speculative decoding and early exit reduce that tax under measurable conditions. Ship them with the same discipline as any production feature: metrics, flags, rollbacks, and clear owners in the serving platform team.