Home / Chapter 3 · How a model writes and what it costs
Last edited · 6 min read
The context window and agents
The model remembers nothing between requests; it knows only what it received in the current one. The context window is the token limit of a single request, shared by input and output. An agent resends the whole growing history every turn, so it pays for that history many times over, and answer quality drops as it grows.
In plain wordsThe model is a consultant with amnesia. Before every meeting they get a folder and know only what is in it. The folder can only be so thick, and the agent adds pages to it at every step.
Move the agent-turn slider. See what fills the window and how fast the cumulative input cost grows
Illustrative numbers. A 200k-token window, 32k of it reserved for the answer and thinking. Each turn adds 6k tokens of tool results and 300 tokens of answer. Rates as in “Prompt caching”: $3 per million input tokens, cache writes at 125%, reads at 10%. Output cost not included.
What counts towards the limit
- Everything in the request: the system prompt, tool definitions (including those from MCP servers), the whole history with tool results, images and PDFs, plus everything the model generates in this turn, thinking included. It is measured in the model’s own tokens, so the same text gives different counts at different providers (see “Tokens”). You can count them before sending with a token-counting endpoint.
- Part of the overhead is fixed and paid every turn: definitions for a few dozen tools easily add up to tens of thousands of tokens before the agent has done anything (see “Tools (function calling)”).
Input limit and output limit
- The window covers input and output together. A separate, much smaller output limit applies per request (
max_tokens): as of September 2026, Claude models with a 1M-token window generate at most 128k per request, and OpenAI’s GPT-6 models have a 1.05M window, at most 922k of it input, and 128k of output. - The API usually rejects input that is too long with an error. Truncating history is a decision made by your application or framework, often silently, starting with the oldest messages.
- When output hits the limit, generation stops mid-sentence:
stop_reason: "max_tokens"in Anthropic (or"model_context_window_exceeded"on Claude 4.5+, when input plus output fill the window),finish_reason: "length"in Chat Completions. JSON is then incomplete, and a reasoning model may never get to the answer. Check the stop reason in code. - Long context can cost more per token: above 200k input tokens Gemini 3.1 Pro, and above 272k GPT-6 models, bill the whole request at 2× for input and cache and 1.5× for output. Claude models with a 1M window have a flat rate (September 2026).
Cost grows with every turn
- In the agent loop (see “The agent loop”), turn t sends everything from earlier turns plus the new result. With a constant increment per turn, cumulative input tokens grow with the square of the number of turns. In the widget, 24 turns already add up to about 2.2M input tokens, even though each request fits in the 200k window.
- Input dominates in agents: Manus reports an average input-to-output token ratio of about 100:1. The main lever on the bill is therefore prompt caching. The causal mask keeps old tokens’ K and V fixed, so as long as the history is only appended to, every turn reads all of it but the newest part from the cache (see “Prompt caching”). Reads cost about 10% of the price, but the curve stays quadratic.
- Latency grows too. Prefill of the uncached part lengthens time to first token, and every decode step reads the whole KV cache, so a long context also slows down generation (see “The generation loop and KV cache”).
- While a tool runs, the model computes nothing, yet the conversation’s KV cache still occupies GPU memory. A busy server can offload the cache to CPU memory or SSD and reload it when the result arrives: faster than recomputing a long history, but not free for time to first token (Glenn Lockwood, September 2026). Through an API you see only the entry lifetime, which at Anthropic counts from the start of the request: generation plus a slow test suite can outlast a 5-minute entry, and the next turn pays for a write and a full prefill of the whole history. For tools that take minutes, the one-hour entry pays off.
Quality drops with length
- Models perform worse on long input well before the window is full. The needle-in-a-haystack test, finding one sentence that matches the question word for word, comes out almost perfect and says little. In RULER (2024), of 17 models claiming at least 32k tokens, only half kept a satisfactory score at 32k.
- The drop is bigger when the question and the answer share no words: in NoLiMa (2025), 11 of 13 models fell below half of their short-context score at 32k tokens. Chroma (2025, 18 models) showed degradation from length alone even on simple tasks, stronger with similar but irrelevant passages.
- Position matters too. In the Lost in the Middle study (2023), models made best use of information at the start and end of the context and markedly worse use of the middle. So put the important instruction and the question at the end, after the long material.
- Window capacity is not a budget to fill. Test your use case at the lengths you actually send, and keep the window short and relevant. How to do that (trimming tool results, compaction, sub-agents, memory outside the window) is covered in “Context engineering and memory”.
Check yourself
What counts towards the context window, and why does a long agent session get expensive and progressively worse?
The model is stateless and only sees the current request. The window caps tokens per request and is shared by the system prompt, tool definitions, history with tool results and everything the model generates, thinking included. The output limit is separate and much smaller. An agent resends the whole history every turn, so cumulative input tokens grow quadratically with turns. Prompt caching lowers the price, not the shape of the curve. Quality degrades with length long before the window is full, so keep context lean and test at realistic lengths.
Po polsku
Model jest bezstanowy i widzi tylko bieżący request. Okno to limit tokenów na request, wspólny dla system promptu, definicji narzędzi, historii z wynikami narzędzi i tego, co model wygeneruje, łącznie z myśleniem. Limit samego wyjścia jest osobny i dużo mniejszy. Agent w każdej turze wysyła całą historię, więc łączne tokeny wejściowe rosną z kwadratem liczby tur. Prompt caching obniża cenę, ale nie kształt krzywej. Jakość spada z długością na długo przed końcem okna, więc kontekst warto trzymać krótki i testować na realnych długościach.
Follow-up questions (4)
- What is context rot?
- Quality degrading with input length before the window runs out. It is stronger when the context holds similar but irrelevant passages and when the answer doesn’t repeat words from the question, so a needle-in-a-haystack test doesn’t reveal it.
- The model has a 1M-token window. Why not put the whole knowledge base in the prompt?
- Because every request pays for the whole million, at some providers at a higher rate above a threshold, time to first token grows, and quality drops with length. For a small, stable set with prompt caching it is reasonable; for a large or changing one, retrieval wins (see “RAG”).
- An agent overflows the window after 40 turns. What do you do?
- First measure what fills the window, usually old tool results and tool definitions. Then trim results at the source, clear old ones, move side tasks to sub-agents, and only as a last resort compact the history (see “Context engineering and memory”).
- Does an API that keeps history server-side lower the cost?
- No. Features such as previous_response_id in OpenAI’s Responses API save on transfer, but the whole history still goes to the model and counts as input. Only prompt caching or a shorter context lowers the cost.