Home / Chapter 4 · Context and knowledge
    Last edited · 10 min read

    Use with AI

    Context engineering and memory

    The model knows only what is in the context of the current call. Context engineering is choosing what goes in: the fewest tokens that are enough for the next step, and from memory outside the window only what is needed right now.

    In plain wordsA hospital shift handover. The night shift doesn’t recount the whole night minute by minute; it leaves a chart: allergies, medication given, what has changed, what is still to do. Whatever isn’t on the chart, the day shift doesn’t know.

    An agent has spent 40 turns adding refunds to a payments module. Choose how it manages context, and see which of six facts survive until the question in turn 41 and whether the agent answers correctly

    input tokens in call 41
    of the 200k window filled
    input tokens across the session
    facts in context

      Illustrative numbers: a 200k-token window with 32k reserved for the answer, an 11k fixed prefix, about 6k tokens per turn, compaction and reset at 150k. The session, facts and agent answers are made up to show typical behaviour. A real agent may lose or keep different facts.

      What competes for window space

      The prompt: the minimum that works

      Techniques for long tasks

      Memory across sessions

      Check yourself

      After an hour of work, an agent forgets what was agreed at the start of the session and keeps getting more expensive. How would you design context and memory management?

      The model knows only what is in the current context, so before every call assemble the smallest set of tokens the next step needs. A stable prefix with instructions and tools goes first, for the cache. The agent fetches large data with tools when it needs it. Clear old tool results and compact history at a threshold, but have the agent write rules and key decisions to a notes file, because a summary loses details. Side tasks go to sub-agents with a clean window. Treat cross-session memory as untrusted data: dated, sourced, expiring and under user control.

      Po polsku

      Model wie tylko to, co jest w bieżącym kontekście, więc przed każdym wywołaniem składaj najmniejszy zestaw tokenów potrzebny do następnego kroku. Stały prefiks z instrukcjami i narzędziami idzie na początek, pod cache. Duże dane agent pobiera narzędziami, gdy ich potrzebuje. Stare wyniki narzędzi czyść, historię przy progu kompaktuj, ale reguły i kluczowe ustalenia agent zapisuje w pliku z notatkami, bo streszczenie gubi szczegóły. Zadania poboczne dostają subagenci z czystym oknem. Pamięć między sesjami traktuj jak niezaufane dane: z datą, źródłem, wygasaniem i kontrolą użytkownika.

      Follow-up questions (5)
      Compaction or a fresh window?
      Compaction preserves the continuity of a long conversation, but the summary loses details and costs an extra call. For code a fresh start is often better: the state lives on disk in notes, tests and git history, and the model can read it back. The Claude docs present this as an alternative to compaction.
      How would you check whether compaction loses something important?
      With evals on long sessions: facts given early, questions about them after compaction, and a comparison with the answer on the full history. Tune the compaction prompt for completeness first, then cut what is unnecessary. Facts that must not be lost, the agent writes to its notes before summarisation happens.
      What goes in the stable prefix, and what do you fetch on demand?
      The prefix holds rules that always apply and short facts needed in most steps. Search finds what resembles the question, and a general rule rarely resembles a specific question. Large, rarely needed data the agent fetches through paths and tools.
      A user says the assistant “remembers” something they never said. What do you check?
      Where the entry came from: a conversation, a document, a web page or a tool result. If it came from foreign content, it is memory poisoning. The fix: write only from what the user says or after their confirmation, with source and date on every entry and a visible notice when something is saved.
      Why not put everything into a million-token window?
      Every call pays for the whole input and waits longer for the first token, and quality drops with length and noise. In the Chroma study, every model tested did markedly better on about 300 tokens of relevant passages than on about 113k tokens in which the same answer was buried in noise. A large window is headroom, not a strategy.

      Sources

      Report an error · Suggest a fix