LLM Serving Schedulers and Continuous Batching Explained
When an AI service handles many language-model requests at once, the inference engine has to decide which prompts and generated tokens use the accelerator at each step. LLM serving schedulers make those decisions, and continuous batching lets them update the active work as requests arrive and finish. Understanding these mechanisms helps infrastructure engineers tune throughput without overlooking response time, memory pressure, or fairness.
What Are LLM Serving Schedulers and Continuous Batching?
An LLM serving scheduler chooses which requests or sequences are ready to run, how much work to submit to the model, and when to admit more work. It coordinates the queue, accelerator execution, and each request’s key-value (KV) cache. The scheduler is part of the serving engine; it does not decide what the model knows or change its weights.
Continuous batching, also called iteration-level or in-flight batching, updates a batch between model execution steps. When a request completes, the engine can free its resources and schedule another request without waiting for every other request in the original group to finish. A scheduler may also combine prompt-processing work with token-generation work, subject to its token and memory budgets.
The vLLM project is one example of an LLM inference and serving engine. Its optimization guide describes scheduling-related controls and chunked prefill, which divides long prompts into smaller pieces that can be scheduled alongside decode work.
Why Serving Schedulers Exist
Language-model requests are uneven. One user may send a short question and request a brief answer; another may submit a long document and generate hundreds of tokens. Requests start at different times, consume different amounts of memory, and finish independently.
A static batch groups requests together and typically holds the group until its slowest member completes. If short requests finish early, their accelerator slots may sit idle while a long response continues. Waiting to build a large batch can also increase the first user’s queue delay. Serving each request alone avoids batch coordination but can leave the accelerator underused.
Continuous batching addresses this mismatch by filling available execution capacity with eligible work as the batch changes. However, a scheduler still has finite compute and KV-cache memory. If it admits too much work, requests may wait longer, cache allocations can run out, and some engines may preempt or recompute sequences. The goal is therefore not maximum concurrency at any cost; it is useful throughput within the service’s latency and memory limits.
How Continuous Batching Works
Text generation has two broad phases. During prefill, the model processes the prompt and creates attention state for its tokens. During decode, it generates output one token at a time, adding new state to the request’s KV cache. Prefill can process many tokens in one operation; decode typically advances each active sequence by one token per iteration.
On each iteration, the scheduler considers active decode sequences and waiting or newly admitted prompts. It selects eligible work within configured limits, executes a model step, then observes which requests finished and what memory is available before scheduling again. The batch is thus refreshed repeatedly instead of being fixed for the entire lifetime of its requests.
| Property | Static batching | Continuous batching |
|---|---|---|
| When membership changes | Usually between batches | Between model iterations |
| Handling a completed request | Other requests may continue occupying the batch | The completed sequence can leave and capacity can be reused |
| Handling new arrivals | Often waits for a later batch | Can be admitted at a subsequent scheduling opportunity |
| Accelerator utilization | Can fall when requests have unequal lengths | Can improve when arrivals and completions are frequent |
| Queueing behavior | Batch formation can add delay | Admission policy and token budgets still determine delay |
| Main operational risk | Idle work slots and batch wait | Contention, memory pressure, and unfairness without controls |
The word continuous describes how the batch is updated, not an absence of queues or synchronization. Model execution still occurs in discrete steps, and requests are not guaranteed immediate admission. In some deployments, streaming generated tokens to the client uses an HTTP event stream; the WHATWG Server-sent events standard defines that transport format, not the inference scheduler.
Components and Scheduling Decisions
The request queue and admission policy hold work that has arrived but has not yet been scheduled. An admission decision can consider queue order, priority, request deadline, tenant, prompt length, or available capacity. A system optimized only for tokens per second may repeatedly favor large or easy batches over an older latency-sensitive request. Throughput and fairness are separate properties; production services often need explicit per-tenant limits, priorities, or queue deadlines.
The token budget caps how much token work can be scheduled in an iteration. In vLLM, max_num_batched_tokens is a relevant serving control. Raising the budget can allow more prompt work into an iteration, but may increase memory demand or delay decode work; a smaller budget can limit per-step work and alter throughput. It is not a maximum prompt length or a promise that every admitted token is processed immediately.
Chunked prefill splits a large prompt into smaller chunks instead of scheduling its entire prefill as one large unit. This gives the scheduler opportunities to mix prompt work with active decode sequences. It can reduce interference between long prompts and ongoing responses, but chunk sizes and scheduling priorities still trade off prompt completion, decode pace, and overall throughput. The right behavior depends on the engine version and workload, so use the engine’s current configuration documentation.
KV-cache allocation determines whether active sequences have memory for their context and generated tokens. Block-based approaches such as PagedAttention manage cache state in units rather than reserving one large contiguous region per sequence; the PagedAttention paper describes this design. Batching and cache allocation are coupled: admitting another sequence consumes resources that may be needed by existing, longer requests. For a deeper explanation of cache memory and reuse, see LLM inference optimization and KV caching.
Backpressure and fairness controls keep overload from turning into unbounded waiting. A service can bound its queue, apply per-user concurrency limits, enforce deadlines, and reject or defer requests when it cannot meet its objectives. Clients should use bounded retries with backoff; immediately retrying every rejected request can amplify load. A scheduler’s prioritization is not by itself an end-to-end fairness guarantee if upstream gateways or downstream worker pools have different queues.
Useful measurements include time to first token (TTFT), inter-token latency (ITL), queue wait, input and output tokens per second, request completion rate, p50 and p95 latency, cache use, and preemptions. Report goodput—the requests or tokens served while meeting latency and quality objectives—alongside raw throughput. Distributed traces can help separate queue delay from model execution; OpenTelemetry for LLM production observability explains how to instrument those paths.
Real-World Use Cases
- Interactive assistants need short and predictable waits while requests of different lengths share accelerators. TTFT and ITL expose different parts of the perceived response delay.
- Document and RAG services can receive long prompts alongside short chat requests. Chunked prefill and admission limits can help control interference, while retrieval quality remains a separate concern.
- Shared inference platforms serve multiple teams or customers. Per-tenant quotas and priorities help stop a single burst from consuming all scheduling capacity.
- Capacity planning uses request traces and representative prompt/output distributions to estimate how concurrency affects goodput, memory use, and latency objectives.
- Separated prefill and decode deployments schedule prompt processing and token generation on distinct worker pools. This can help when the phases have different resource needs, but requires KV-state transfer and additional coordination. Speculative decoding and disaggregated LLM inference covers those trade-offs.
Getting Started: Measure Continuous Batching with vLLM
Run serving benchmarks on a supported Linux GPU environment. Follow the current vLLM GPU installation guide because supported accelerators, drivers, and installation options vary. With uv installed, a basic setup is:
uv pip install vllm --torch-backend=auto
vllm serve Qwen/Qwen2.5-1.5B-Instruct \
--max-model-len 4096 \
--max-num-batched-tokens 2048 \
--max-num-seqs 32
These example limits are starting points, not universal tuning values. max-num-batched-tokens bounds token work scheduled in an iteration, while max-num-seqs limits scheduled sequences. Choose values supported by the model, engine version, and available memory; the vLLM optimization guide explains relevant behavior, including preemption and chunked prefill.
Confirm the server responds and make a small streaming request:
curl http://localhost:8000/v1/models
curl -N http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"Qwen/Qwen2.5-1.5B-Instruct","messages":[{"role":"user","content":"Explain continuous batching in two sentences."}],"max_tokens":64,"stream":true}'
To compare configurations, replay the same representative request mix at controlled arrival rates. Include short and long prompts, different output lengths, and concurrent users; hold the model, hardware, and workload constant. Measure TTFT, ITL, queue time, p95 completion latency, successful requests per second, tokens per second, GPU memory, and preemptions. Inspect whether requests meet their latency objectives, not just whether aggregate token throughput rises.
Change one limit at a time. If memory pressure causes preemption, reduce concurrency or the scheduled token budget and compare again. If prompts delay active generations, test the documented chunked-prefill behavior and observe ITL. If queue times grow while the GPU is saturated, increase capacity or enforce admission controls rather than allowing an unlimited queue. A local smoke test demonstrates that the server works; it does not substitute for a load test or a quality evaluation.
Common Misconceptions
- “Continuous batching removes queues.” It can admit work between iterations, but when demand exceeds available compute or memory, requests still wait or are rejected.
- “A larger batch is always faster.” Larger work groups can improve utilization, but can also increase queueing, memory pressure, preemption, or latency for active requests.
- “Streaming and continuous batching mean the same thing.” Streaming sends response fragments to a client; continuous batching changes how inference work is grouped and scheduled. A service may use either without the other.
- “Higher tokens per second proves better service.” Aggregate throughput can rise while individual requests violate latency targets or tenants receive unequal service. Measure goodput and fairness alongside throughput.
Related Articles
- LLM Inference Optimization and KV Caching Explained covers cache memory and reuse.
- Speculative Decoding and Disaggregated LLM Inference explains draft verification and prefill/decode separation.
- Inference-Time Scaling and Test-Time Compute Explained discusses request-level compute and its latency trade-offs.
- OpenTelemetry for LLM Production Observability covers tracing and measurement.
Changelog and Last Updated
Last updated: October 9. Initial publication.

