Speculative Decoding and Disaggregated LLM Inference

Updated on
9 min read

As language-model applications add longer prompts, larger models, and more concurrent users, serving teams look beyond model size to the way each request uses accelerators. Speculative decoding and disaggregated LLM inference are two system-level approaches to improve serving behavior: one proposes multiple output tokens for verification, while the other assigns prompt processing and token generation to separate worker pools. This guide explains how each works, what to measure, and when the added complexity is worthwhile.

What Are Speculative Decoding and Disaggregated Inference?

In speculative decoding, a smaller or cheaper process proposes several likely next tokens. The larger target model checks those candidates together and accepts the valid prefix; if a candidate does not match, normal generation resumes from the correction. With an appropriate verification algorithm, the output distribution remains that of the target model rather than the draft model. The vLLM speculative decoding guide describes supported approaches and their configuration.

Disaggregated inference separates the two main phases of autoregressive generation. A prefill worker processes the input prompt in parallel and creates the attention key-value (KV) state. A decode worker then generates output tokens sequentially, using that state. The KV state must be transferred or made accessible between the workers. The DistServe research paper studies this separation as a way to tune resources for each phase independently.

These techniques address different bottlenecks and can be combined. Neither changes the model’s learned weights, and neither guarantees a faster request by itself.

The Problem They Solve

An autoregressive model cannot ordinarily emit its next token until it has computed the current one. Even when a GPU runs the model quickly, each generation step waits for the previous step. If decoding is limited by memory traffic or the target model’s per-token latency, speculative decoding can use otherwise idle compute to verify multiple draft tokens in a single target-model pass. It helps most when the draft is cheap and its proposals are often accepted.

The prompt phase has a different shape: many input tokens can be processed together, which can create a burst of compute and memory demand. Mixing long prompts with token-by-token decoding on the same accelerators can make it harder to meet both time-to-first-token and inter-token-latency targets. Disaggregation lets operators allocate and schedule prompt-heavy prefill separately from decode-heavy generation.

Both changes introduce costs. Drafting consumes compute and memory, while separated workers need coordination and a path for KV data. The relevant goal is goodput: the number of useful requests or tokens served while staying within latency and quality objectives, not just a higher raw token count.

How the Architectures Work

Speculative decoding begins with a draft model or another proposal mechanism, such as a reusable prompt lookup. It produces a short candidate sequence. The target model evaluates the candidates in parallel, accepting a prefix and sampling a correction or continuing normally at the first rejection. The process repeats until the response is complete. A longer proposal can reduce target-model steps when acceptance is high, but it also spends more time on proposals that may be discarded. The vLLM project documents speculative decoding as one option in a broader serving engine; actual performance depends on the model pair, workload, hardware, and implementation.

Disaggregated prefill and decode route a request to a prefill worker, which processes the prompt and produces KV state. A scheduler then hands the request and its state to a decode worker. That worker produces tokens and may stream them back through an API gateway. KV transfer can use a network, shared memory, or other supported transport. The vLLM disaggregated prefilling guide and NVIDIA Dynamo documentation describe serving designs that separate these phases.

The API boundary is not the same thing as the inference architecture: two workers can communicate behind one client-facing endpoint. Standards such as HTTP semantics in RFC 9110 describe the request and response layer, not how an inference server schedules GPU work or moves KV state.

Dimension Speculative decoding Disaggregated prefill and decode
Primary bottleneck Sequential target-model token steps Contention between prompt processing and generation
Main components Target model plus a draft or proposal mechanism Separate prefill and decode workers, scheduler, and KV-transfer path
Core operation Verify several proposed tokens together Process prompt and generated tokens on different worker pools
Latency measure to watch Inter-token latency and end-to-end duration Time to first token, inter-token latency, and queueing
Potential benefit Fewer target-model decoding steps Independent resource allocation and less phase interference
Added cost Draft computation, memory, and proposal verification KV movement, routing, synchronization, and extra capacity
Good fit when Draft tokens are accepted often and decoding is the bottleneck Prompt and decode demand have different resource profiles
Relationship Can run within a decode worker Can host decode workers that also use speculative decoding

Components and Key Concepts

Draft quality and acceptance rate determine whether speculative decoding can save work. A draft that predicts likely continuations well can increase the number of accepted tokens per target-model pass. A draft that is too large, slow, or poorly matched to the target can cost more than it saves. Some implementations use a separate model; others propose from n-grams or matching prompt text.

KV state and transport are central to disaggregated serving. KV state grows with prompt and generated context, so transferring it can consume substantial bandwidth and add latency. A remote decode worker is useful only when the benefit of specialization outweighs the state-transfer and coordination costs. Reusing or retaining KV data is a related but separate optimization; see LLM inference optimization and KV caching.

Scheduling and capacity matter in both designs. A shared GPU pool may leave too little memory for a draft model or allow long prefill jobs to delay decode work. Separate pools require demand-aware routing: too few prefill workers can raise prompt queues, while too few decode workers can slow token emission.

Measure the entire request under representative prompt lengths, output lengths, concurrency, and arrival rates. Track time to first token (TTFT), inter-token latency (ITL), p50 and p95 request duration, input and output throughput, accelerator memory, KV transfer volume, and draft acceptance. Include output quality and error rates. For end-to-end request traces across retrieval and model calls, see OpenTelemetry for LLM production observability.

Real-World Use Cases

  • Interactive assistants: Speculative decoding may improve perceived response speed when users receive a streamed answer and the target model is decode-bound. Measure ITL as well as total duration; a quicker first token does not guarantee a faster full response.
  • Long-context or RAG services: Prefill can dominate the wait for the first token when requests contain long retrieved context. A separate prefill pool can isolate that workload from ongoing decode, although transferring a large KV state may offset the gain. Hybrid search for RAG covers the retrieval stage that happens before model inference.
  • Mixed request workloads: A service receiving both short chat prompts and long document-analysis prompts may benefit from separate phase capacity if their demand patterns differ consistently.
  • Large deployments with varied accelerators: Operators may place prompt processing and decoding on hardware suited to their respective compute and memory profiles. This is useful only when interconnect bandwidth, scheduling, and utilization support the split.

For a foundation in the self-attention computations behind prompt processing and generation, read Transformer architecture deep dive.

Getting Started: Test Speculative Decoding with vLLM

Start in a supported Linux GPU environment and follow the vLLM installation guide; model and hardware compatibility varies. With uv installed, install vLLM:

uv pip install vllm --torch-backend=auto

First run the target model without speculation to establish a baseline:

vllm serve Qwen/Qwen2.5-1.5B-Instruct --max-model-len 4096

Stop that server before starting the speculative configuration. This example uses a smaller model from the same family as the draft; check current vLLM compatibility requirements before using a different pair:

vllm serve Qwen/Qwen2.5-1.5B-Instruct \
  --speculative-config '{"model":"Qwen/Qwen2.5-0.5B-Instruct","num_speculative_tokens":5}' \
  --max-model-len 4096

The server downloads the selected models if they are not already available. Confirm it responds, then send a request:

curl http://localhost:8000/v1/models

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen/Qwen2.5-1.5B-Instruct","messages":[{"role":"user","content":"Explain speculative decoding briefly."}],"max_tokens":96}'

Compare baseline and speculative runs with the same hardware, model, prompts, generation limits, and request load. Repeat enough requests to account for warm-up and variability. Record TTFT, ITL, throughput, memory use, and output quality; inspect acceptance behavior if the serving engine exposes it. Do not infer a win from one short prompt.

For disaggregation, begin with a separate prefill/decode deployment using an inference framework that documents the required scheduler and KV-transfer mechanism. The NVIDIA Dynamo deployment documentation provides architecture and configuration guidance; a production setup also needs capacity planning, network and failure testing, and telemetry for transfer time and worker queues. Keep a co-located baseline so the extra network and operational costs can be compared.

Common Misconceptions

  • “Speculative decoding changes the answer to match the draft.” A correct verification procedure preserves the target model’s output distribution. The draft proposes candidates; it does not replace the target.
  • “A small draft always makes generation faster.” The draft adds work. If it is slow, acceptance is low, or decoding is not the bottleneck, latency or throughput can get worse.
  • “Disaggregation eliminates prefill or decode work.” It changes where phases run and how resources are allocated. Both phases still execute, and KV state must still be stored and transferred or shared.
  • “These techniques are competing alternatives.” They address different parts of inference and may be combined, but each should be evaluated against its own baseline and operational cost.

Changelog and Last Updated

Last updated: October 4. Initial publication.

TBO Editorial

About the Author

TBO Editorial writes about the latest updates about products and services related to Technology, Business, Finance & Lifestyle. Do get in touch if you want to share any useful article with our community.