Home / Chapter 3 · How a model writes and what it costs
    Last edited · 6 min read

    Use with AI

    Prompt caching

    Prompt caching is the KV cache kept by the provider between requests. A request that starts exactly like a recent one skips prefill for that part: it pays a fraction of the input price and gets its first token sooner.

    In plain wordsA waiter who knows the regulars doesn’t ask again about your allergies and favourite table. But give a different name and they start from scratch. And they remember you for only a few minutes after your last visit.

    Pick a prompt layout. See how much of the second request hits the cache and what it costs

    Request 1

    Request 2, a moment later

    cache writehitrecomputed and written
    of request 2 tokens come from the cache
    request 1 input cost: write only
    request 2 input cost: read plus writing the new part
    cheaper than without caching

    Example rates: $3 per million input tokens, cache reads at 10% of that price, writes at 125%. This is how Anthropic prices a 5-minute entry, and OpenAI from GPT-5.6 onwards (September 2026). With older models and other providers, writes often carry no surcharge.

    How it works

    Provider terms

    How to lay out the prompt

    Check yourself

    How does prompt caching work, and how do you structure a prompt for it?

    It is the KV cache kept by the provider across requests. It only matches an identical prefix, because a token’s keys and values depend on everything before it: the first changed token invalidates the rest. So stable parts such as tools, the system prompt and large documents go first, history is append-only, and volatile data goes last. Reads usually cost 10% of the input price, writes can cost more than plain input, and entries live for minutes. The payoff is cheaper input and faster time to first token. The hit rate comes from the usage fields.

    Po polsku

    To KV cache zachowany u dostawcy między requestami. Działa tylko na identyczny prefiks, bo K i V tokena zależą od wszystkiego przed nim: pierwszy zmieniony token unieważnia resztę. Dlatego stałe części, czyli narzędzia, system prompt i duże dokumenty, idą na początek, historię się tylko dopisuje, a zmienne dane trafiają na koniec. Odczyt kosztuje zwykle 10% ceny wejścia, zapis bywa droższy od zwykłego wejścia, a wpis żyje minuty. Zysk to tańsze wejście i krótszy czas do pierwszego tokena. Hit rate odczytuje się z pól usage.

    Follow-up questions (4)
    The agent bill is high and the hit rate is low. What do you check?
    Diff consecutive requests token by token and look for the first difference: a date or ID at the start, non-deterministic serialisation of tools and JSON, edited history, a change of model or thinking level. Then check whether the gaps between turns exceed the entry lifetime and whether the prefix meets the minimum length.
    How is prompt caching different from a semantic cache?
    Prompt caching stores the computation for an identical prefix, and the model still generates a new answer, so quality doesn’t change. A semantic cache returns an old answer to a similar question: it saves the whole call but may return the answer to a different question.
    When does caching not pay off?
    When the prefix rarely repeats within the entry’s lifetime. With a write surcharge, every entry that never gets a hit costs more than a request without caching.
    Can the cache leak another user’s data?
    Large providers don’t share the cache across organisations, but within your application users with the same prefix hit the same entry. A hit shows up as a shorter response time, so someone can test whether another user recently sent a given text. A 2025 audit found cache shared globally, across organisations, at 7 of 17 API providers, and at least five changed that only after disclosure. Isolate the cache per customer (a separate workspace at Anthropic; at OpenAI, from GPT-5.6, a separate prompt_cache_key, while on older models the key only steers routing) or keep sensitive data out of the shared prefix.

    Sources

    Report an error · Suggest a fix