Home / Chapter 4 · Context and knowledge
Last edited · 10 min read
Context engineering and memory
The model knows only what is in the context of the current call. Context engineering is choosing what goes in: the fewest tokens that are enough for the next step, and from memory outside the window only what is needed right now.
In plain wordsA hospital shift handover. The night shift doesn’t recount the whole night minute by minute; it leaves a chart: allergies, medication given, what has changed, what is still to do. Whatever isn’t on the chart, the day shift doesn’t know.
An agent has spent 40 turns adding refunds to a payments module. Choose how it manages context, and see which of six facts survive until the question in turn 41 and whether the agent answers correctly
Illustrative numbers: a 200k-token window with 32k reserved for the answer, an 11k fixed prefix, about 6k tokens per turn, compaction and reset at 150k. The session, facts and agent answers are made up to show typical behaviour. A real agent may lose or keep different facts.
What competes for window space
- A single call includes the system prompt, tool definitions, examples, retrieved documents, conversation history, tool results and the model’s reasoning (thinking). In agents, tool results grow fastest. Manus reports an average of about 100 input tokens per output token. The window limit and how cost grows with every turn are covered in “The context window and agents”.
- Quality drops before the window runs out. This is context rot. Chroma (July 2025) tested 18 models: results got worse with input length even on simple tasks, and passages on the same topic that don’t answer the question did extra damage. On LongMemEval, every model did markedly better on the version with only the relevant passages (about 300 tokens) than on the full one (about 113k), even though the answer was in both. Anthropic explains this as an attention budget: n tokens make n² pairs, so attention is spread across ever more candidates (see “Attention”). Fewest tokens doesn’t mean short: a missing fact hurts as much as noise.
The prompt: the minimum that works
- Task, reason, constraints and output format. A reason works better than a prohibition. Instead of “NEVER use ellipses”, say that the text will be read by a speech synthesiser that can’t pronounce an ellipsis. From the reason the model also infers cases you didn’t list. Say what to do rather than what not to do. The test from the Claude docs: would a colleague with no context be able to follow this instruction?
- Examples steer format and tone more effectively than a description. The Claude docs recommend 3–5 varied examples in
<example>tags. Anthropic advises a few typical ones rather than a list of every edge case. The model copies examples that are too similar together with their accidental features, such as always the same length. - Structure comes from sections or XML tags: instructions, context, examples and input data kept apart, so the model doesn’t confuse instructions with data. Long documents go at the top, the question at the end. In Anthropic’s tests this gave up to 30% better answers with multiple documents.
Techniques for long tasks
- Stable prefix first: system prompt, tool definitions, fixed documents. Variable content goes last, history is append-only, and serialisation is deterministic (the same JSON key order). Then consecutive calls hit the cache; see “Prompt caching”. Any change in the middle of the context, such as compaction, clearing or a new tool, invalidates the cache from that point. So edit rarely and in large chunks. For this reason Manus doesn’t remove tools mid-task; it masks them out during decoding instead.
- Just-in-time retrieval instead of loading everything up front. The agent keeps lightweight pointers (file paths, URLs, queries) and reads the content with tools when it needs it:
grep,head, a database query. Claude Code combines both:CLAUDE.mdgoes into the context at the start, and the agent reads files as it goes. Agent Skills do the same with instructions: the prefix holds only each skill’s name and description, about 100 tokens, and the full instructions load when a task matches. It is the same fix as tool search for bloated tool definitions (see “MCP”); more in “The coding-agent harness”. The price: more turns and slower than ready-made results from an index (see “RAG”). Load rules that always apply up front. Search finds what resembles the question, and “nothing on production without approval” resembles none. - Compaction: at a threshold the model summarises the history, and work continues from the summary and the latest turns. Claude Code keeps architectural decisions, unresolved bugs and implementation details, drops repeated tool results and adds the 5 most recently read files. The summary loses details that looked unimportant at the time, and producing it costs a call that reads the whole history. As of September 2026 the Claude API (beta) and OpenAI’s Responses API can compact on the server, at a token threshold or on request. Claude returns a readable summary and accepts your own summarisation prompt; OpenAI’s compaction item is encrypted, so only evals show what it lost.
- Clearing old tool results is the gentlest form. A raw result from 30 turns ago is rarely needed verbatim. You replace it with a placeholder while the call itself stays, so the agent knows what it has already checked and can fetch it again. In the Claude API this is context editing: past a threshold (100k input tokens by default) it clears older results and keeps the 3 most recent. On Claude Fable 5.1 and Opus 5.5, if you send thinking blocks back, leave clearing to the API. Editing an earlier tool result in your own code, or compacting while keeping recent turns with their thinking, fails the prefix check for every later thinking block: a 400 by default for accounts created from 31 August 2026. One summary that replaces the whole history is fine.
- Notes outside the window: the agent maintains a file itself, such as
NOTES.md, a to-do list orprogress.txt, and reads it after a context reset. Manus, averaging about 50 tool calls per task, keeps rewritingtodo.mdso the goal sits at the end of the context instead of getting lost in the middle. For coding work the Claude docs suggest sometimes starting from a clean window instead of compacting: the model rebuilds its state from notes, tests and git history. Notes contain only what the agent judged important, and they go stale too. - A sub-agent gets a clean window for a side task, such as searching a repository. It uses tens of thousands of tokens and returns a summary, according to Anthropic usually 1–2k tokens. The main context stays short, but the sub-agent doesn’t see the main thread’s decisions, so the brief it gets must include them. More in “Multiple agents”.
Memory across sessions
- Long-term memory is a store outside the model that comes back into context in later sessions. It holds facts and preferences (works in TypeScript, prefers short answers), episodes (how a similar ticket was resolved, what didn’t work) and rules (corrected instructions). It takes one of two forms: a single profile rewritten as a whole, which is simple but easily loses something on update, or a collection of small entries, which loses less but is harder to search, correct and delete.
- Memory comes back into context in two ways. A small profile always sits in the prefix: simple, but you pay for it on every call. A larger store is searched by a tool when the agent needs it: it scales, but may miss. The memory tool in the Claude API runs client-side. The model requests file operations in a
/memoriesdirectory, your code carries them out on your storage, and the API adds an instruction to always check that directory before starting work. Writing during work is visible immediately but slows things down. Writing in the background after the session doesn’t slow anything down but takes effect with a delay. - Memory goes stale. An entry saying “the project uses MySQL” after a migration to Postgres does more harm than no entry, because the model trusts its notes. Keep a date and source with each entry, let the newer one win, and let unused entries expire. The memory tool docs add a file size limit, stripping sensitive data before saving, and protection against paths like
/memories/../../. - Memory poisoning: foreign content (a web page, document or email) tells the model to save a false entry, which then takes effect in every session. In 2024 Johann Rehberger showed false memories being planted in ChatGPT via documents, images and browsed pages, and then an entry that sent all future conversations to the attacker. According to the author, OpenAI blocked that exfiltration channel in September 2024, but not the memory writes themselves. Treat writing to memory as a privileged action: only from what the user says or with their consent, and with a visible notice. See “Prompt injection”.
- Users must be able to see, correct and delete their memory, and to chat without it. One user’s memory must never reach another user’s context: a separate space per user and project, and when an account is deleted, its memory goes too.
Check yourself
After an hour of work, an agent forgets what was agreed at the start of the session and keeps getting more expensive. How would you design context and memory management?
The model knows only what is in the current context, so before every call assemble the smallest set of tokens the next step needs. A stable prefix with instructions and tools goes first, for the cache. The agent fetches large data with tools when it needs it. Clear old tool results and compact history at a threshold, but have the agent write rules and key decisions to a notes file, because a summary loses details. Side tasks go to sub-agents with a clean window. Treat cross-session memory as untrusted data: dated, sourced, expiring and under user control.
Po polsku
Model wie tylko to, co jest w bieżącym kontekście, więc przed każdym wywołaniem składaj najmniejszy zestaw tokenów potrzebny do następnego kroku. Stały prefiks z instrukcjami i narzędziami idzie na początek, pod cache. Duże dane agent pobiera narzędziami, gdy ich potrzebuje. Stare wyniki narzędzi czyść, historię przy progu kompaktuj, ale reguły i kluczowe ustalenia agent zapisuje w pliku z notatkami, bo streszczenie gubi szczegóły. Zadania poboczne dostają subagenci z czystym oknem. Pamięć między sesjami traktuj jak niezaufane dane: z datą, źródłem, wygasaniem i kontrolą użytkownika.
Follow-up questions (5)
- Compaction or a fresh window?
- Compaction preserves the continuity of a long conversation, but the summary loses details and costs an extra call. For code a fresh start is often better: the state lives on disk in notes, tests and git history, and the model can read it back. The Claude docs present this as an alternative to compaction.
- How would you check whether compaction loses something important?
- With evals on long sessions: facts given early, questions about them after compaction, and a comparison with the answer on the full history. Tune the compaction prompt for completeness first, then cut what is unnecessary. Facts that must not be lost, the agent writes to its notes before summarisation happens.
- What goes in the stable prefix, and what do you fetch on demand?
- The prefix holds rules that always apply and short facts needed in most steps. Search finds what resembles the question, and a general rule rarely resembles a specific question. Large, rarely needed data the agent fetches through paths and tools.
- A user says the assistant “remembers” something they never said. What do you check?
- Where the entry came from: a conversation, a document, a web page or a tool result. If it came from foreign content, it is memory poisoning. The fix: write only from what the user says or after their confirmation, with source and date on every entry and a visible notice when something is saved.
- Why not put everything into a million-token window?
- Every call pays for the whole input and waits longer for the first token, and quality drops with length and noise. In the Chroma study, every model tested did markedly better on about 300 tokens of relevant passages than on about 113k tokens in which the same answer was buried in noise. A large window is headroom, not a strategy.