Home / Chapter 7 · Production and serving
    Last edited · 5 min read

    Use with AI

    Continuous batching and PagedAttention

    A server groups conversations into a batch so that one read of the weights serves many users at once. Continuous batching swaps conversations in and out of the batch at every step, and PagedAttention hands out KV cache memory in small blocks, so the same GPU fits several times more of them.

    In plain wordsA bus that doesn’t wait for everyone to reach the terminus: whoever gets off frees a seat for the next passenger at the very next stop. And nobody reserves a seat for the whole route in advance; you take one when you board.

    Play both modes and compare how many cells sit empty and at which step the last conversation finishes

    conversation generates a tokenslot sits empty
    conversations finished
    empty cells (slot × step) so far
    steps to serve all eight

    Eight conversations (A–H), 2 to 9 tokens long. Rows s1–s4 are the four slots in the batch; columns are successive steps. Illustrative: it leaves out the prefill a new conversation must run before its first token.

    How the scheduler works

    Switch the memory allocation scheme and count how many conversations fit in the same 32 blocks

    block used by a conversationreserved but empty
    conversations fit in memory
    of memory reserved but unused

    Illustrative: the reservation assumes a maximum of 8 blocks per conversation. In a real server the tail of each conversation’s last block also stays empty.

    PagedAttention: KV cache memory

    What it means in production

    Check yourself

    How does an LLM server handle many users at once, and what do continuous batching and PagedAttention give you?

    A batch shares the cost of reading the weights, so a server wants it as large as possible. Continuous batching decides batch membership at every decode step rather than per request: a finished conversation frees its slot immediately and a new one joins on the next step, so the GPU never waits for the longest answer. PagedAttention allocates the KV cache in blocks on demand instead of reserving the maximum up front: memory waste drops from 60–80% to a few per cent, more conversations fit, and shared prefixes are stored once. The price is preemption when memory runs out.

    Po polsku

    Batch dzieli koszt czytania wag, więc serwer chce go mieć jak największy. Continuous batching ustala skład batcha co krok decode, a nie co żądanie: skończona rozmowa od razu zwalnia miejsce, nowa dołącza w następnym kroku, więc karta nie czeka na najdłuższą odpowiedź. PagedAttention przydziela KV cache blokami na żądanie zamiast rezerwować maksimum z góry: strata pamięci spada z 60–80% do kilku procent, mieści się więcej rozmów, a wspólne prefiksy leżą raz. Ceną są wywłaszczenia, gdy pamięć się skończy.

    Follow-up questions (4)
    What limits the batch size?
    KV cache memory for all conversations and the latency budget between tokens. With long contexts memory runs out first; with short ones, the latency budget.
    Users report rare stalls of a few seconds in the token stream. What do you check?
    Preemptions caused by running out of KV cache memory (a counter in the server metrics), and long prefills of new requests without chunked prefill. Fixes: a lower limit on concurrent conversations, more memory for the cache, or separate replicas for very long prompts.
    What does PagedAttention give you beyond tighter packing?
    Block sharing with copy-on-write. A prefix common to many conversations is stored once, and several samples of the same answer (n > 1, beam search) share the prompt’s blocks. The vLLM authors report up to 55% less memory with such sampling.
    Why does the same prompt at temperature 0 sometimes produce different text?
    A kernel’s output depends on the batch size, because the order of floating-point additions changes. Batch size depends on server load, so small differences in the logits sometimes change the chosen token. Thinking Machines (2025) identifies this as the main cause of nondeterminism in inference endpoints. The fix is batch-invariant kernels: vLLM has a Batch Invariance mode (beta as of September 2026) whose output doesn’t depend on batch size or request order, which can cost some speed. It is useful for evals, debugging and RL rollouts.

    Sources

    Report an error · Suggest a fix