## Continuous batching

*Production and serving*

*Last edited: 28 September 2026*

A server groups conversations into a batch so that one read of the weights serves many users at once. Continuous batching swaps conversations in and out of the batch at every step, and PagedAttention hands out KV cache memory in small blocks, so the same GPU fits several times more of them.

**In plain words:** A bus that doesn’t wait for everyone to reach the terminus: whoever gets off frees a seat for the next passenger at the very next stop. And nobody reserves a seat for the whole route in advance; you take one when you board.

*Interactive widget on the page: Play both modes and compare how many cells sit empty and at which step the last conversation finishes.*

### How the scheduler works

- A static batch fixes its membership once: everyone waits for the longest answer, and new requests wait for the whole group to finish. You don’t know answer lengths in advance, so empty slots are the rule, not the exception.
- Continuous batching (iteration-level scheduling in the literature, from the Orca system, OSDI 2022) decides the batch membership at every decode step. A finished conversation frees its slot immediately and a waiting one joins on the next step. It works because each step is a separate forward pass, and the layers other than attention don’t depend on sequence length.
- A new conversation starts with the prefill of its prompt, which briefly takes the GPU away from conversations that are mid-generation. Chunked prefill cuts a long prompt into chunks that fit the per-step token budget. vLLM V1 enables it by default: each step it schedules the decoding conversations first and fills the remaining budget with prefill.

*Interactive widget on the page: Switch the memory allocation scheme and count how many conversations fit in the same 32 blocks.*

### PagedAttention: KV cache memory

- Reserving a contiguous region for the maximum answer length wastes memory on headroom and fragmentation. The vLLM authors measured 60–80% of KV cache memory wasted in the systems of the time.
- PagedAttention splits the KV cache into blocks of a dozen or so tokens (usually 16), allocated as generation proceeds and addressed through a block table, like pages of virtual memory in an operating system. Only the tail of the last block stays empty, and waste drops below 4%. More conversations in memory means a bigger batch and a lower cost per token.
- Blocks can be shared with copy-on-write. A common prefix, such as a system prompt or several samples of the same answer, is stored in memory once. Server-side prefix caching builds on this (see “Prompt caching”). With several replicas it pays off only if the router sends requests with the same prefix to the replica that already holds it (prefix-aware routing, e.g. in llm-d or NVIDIA Dynamo).
- The price of on-demand allocation is that memory can run out mid-generation. The server then preempts a conversation and later recomputes its KV cache from scratch. In vLLM V1 this is the default, rather than swapping the cache out to CPU memory. The user sees it as a sudden stall in the stream.

### What it means in production

- Batch size is limited by KV cache memory (number of conversations × context length) and by the latency budget between tokens. Server settings such as `max_num_seqs` and `max_num_batched_tokens` in vLLM set this trade-off directly.
- Throughput comes at the cost of latency. Under load, time to first token grows because of the queue and prefills, and time between tokens grows because of larger batches. Measure the two separately, at p50 and p99 (see “LLMs in production”).
- Providers’ batch APIs sell a loose SLO: at Anthropic it is 50% cheaper, most batches finish within an hour, and the limit is 24 hours. The provider can use this traffic to fill free batch slots and traffic troughs. For evals and bulk processing it is the simplest saving.

### Check yourself

**Question:** How does an LLM server handle many users at once, and what do continuous batching and PagedAttention give you?

**Short answer:** A batch shares the cost of reading the weights, so a server wants it as large as possible. Continuous batching decides batch membership at every decode step rather than per request: a finished conversation frees its slot immediately and a new one joins on the next step, so the GPU never waits for the longest answer. PagedAttention allocates the KV cache in blocks on demand instead of reserving the maximum up front: memory waste drops from 60–80% to a few per cent, more conversations fit, and shared prefixes are stored once. The price is preemption when memory runs out.

### Follow-up questions

- **What limits the batch size?** KV cache memory for all conversations and the latency budget between tokens. With long contexts memory runs out first; with short ones, the latency budget.
- **Users report rare stalls of a few seconds in the token stream. What do you check?** Preemptions caused by running out of KV cache memory (a counter in the server metrics), and long prefills of new requests without chunked prefill. Fixes: a lower limit on concurrent conversations, more memory for the cache, or separate replicas for very long prompts.
- **What does PagedAttention give you beyond tighter packing?** Block sharing with copy-on-write. A prefix common to many conversations is stored once, and several samples of the same answer (n > 1, beam search) share the prompt’s blocks. The vLLM authors report up to 55% less memory with such sampling.
- **Why does the same prompt at temperature 0 sometimes produce different text?** A kernel’s output depends on the batch size, because the order of floating-point additions changes. Batch size depends on server load, so small differences in the logits sometimes change the chosen token. Thinking Machines (2025) identifies this as the main cause of nondeterminism in inference endpoints. The fix is batch-invariant kernels: vLLM has a Batch Invariance mode (beta as of September 2026) whose output doesn’t depend on batch size or request order, which can cost some speed. It is useful for evals, debugging and RL rollouts.

### Sources

- [Kwon et al.: Efficient Memory Management for LLM Serving with PagedAttention (SOSP 2023)](https://arxiv.org/abs/2309.06180)
- [vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention](https://vllm.ai/blog/2023-06-20-vllm)
- [vLLM: Optimization and Tuning (preemption, chunked prefill)](https://docs.vllm.ai/en/latest/configuration/optimization.html)
- [Agrawal et al.: Sarathi-Serve, chunked prefill (OSDI 2024)](https://arxiv.org/abs/2403.02310)

Interactive page: https://howaiworks.dev/continuous-batching/
