Home / Chapter 5 · Agents
Last edited · 10 min read
The coding-agent harness
A coding agent is a model plus a harness: everything around the model that turns it into a working agent. The model proposes the next step. The harness decides what the model sees, what actually runs and what can be undone.
In plain wordsA skilled contractor on a building site. The skill is theirs, but the site decides the rest: which tools are laid out, what the job sheet and site rules say, which rooms are locked, who has to sign off before a wall comes down, and whether there is a spirit level to check the work. The same contractor does very different work on a well-run site and on a chaotic one.
Replay one task in a coding agent: partial refunds are 1 cent short. Switch off one part of the harness, press ▶ and see what changes
Illustrative numbers: a 200k window, a 12k system prompt with built-in tools, 3k of project memory, about 5k tokens of history per turn, compaction at 80% of the window, 50k for all MCP definitions loaded up front. The subagent’s own tokens are not counted. The repository, commands and model behaviour are made up to show typical failures.
What the harness is made of
- The loop from “The agent loop” with a few general tools: read a file, edit by replacing a string, run a shell command, search with grep and glob. The harness validates every call, runs it and clips long output: Claude Code writes an MCP result above 25k tokens to a file and gives the model the path (Claude Code and Codex details on this page are as of September 2026).
- Instructions. The system prompt describes the tools and working rules, and a project memory file is loaded at the start of every session. AGENTS.md is an open format read by Codex, Cursor, Copilot, Jules and others. Claude Code reads CLAUDE.md, or AGENTS.md when there is no CLAUDE.md. The file is context, not configuration: the model usually follows it, but nothing forces it to.
- Planning. The model writes a todo list and ticks it off, and the harness puts the list back at the end of the context whenever it changes, so the goal doesn’t sink into the middle of a long history. In plan mode the agent may only read and run read-only commands until you accept the plan.
- Context management, covered in “Context engineering and memory”: clearing old tool results, compaction at a threshold, and subagents that search in their own window and return a summary. Two mechanisms keep material out of the window until it is needed. Tool search puts only tool names in the prefix and loads full definitions on demand, the Claude Code default. Agent Skills, an open format, are folders with a SKILL.md file: only the name and description sit in the prefix, and the body loads when the task matches (progressive disclosure).
- Permissions decide which calls wait for a human. In Claude Code, reads and read-only commands such as
ls,grepandgit statusrun freely, while other shell commands and edits ask by default. Modes range from plan, through accept-edits and auto (a classifier reviews each action), to bypass. A sandbox is enforced by the operating system (Seatbelt on macOS, bubblewrap on Linux): writes only in the working directory, network only through a proxy with a domain allowlist. Codex in its default workspace-write mode has the network off. Both layers are needed: without network isolation a hijacked agent can send out your SSH keys, and without filesystem isolation it can plant something that opens the network later. - Hooks are your commands at fixed points in the loop: before a tool call (they can block it), after an edit (a formatter), when the agent wants to finish (tests). They run every time, unlike an instruction the model may skip. In Claude Code a hook that denies a call blocks it even in bypass mode. MCP servers add outside tools (see “MCP”).
- Checkpoints. Claude Code snapshots files before each of your prompts, and
/rewindrestores code, conversation or both. It doesn’t track files changed by shell commands such asrmormv, nor effects outside your machine, so git stays the real undo.
Same model, different harness
- SWE-bench and Terminal-Bench measure a model and a harness together. Anthropic noted in January 2025 that scores vary significantly with the scaffold even for the same model. LangChain (February 2026, a report on its own product) kept GPT-5.2-Codex fixed and changed only the system prompt, tools and hooks, including a checklist that forces verification before the agent finishes and loop detection: Terminal-Bench 2.0 rose from 52.8% to 66.5%. It works the other way too: in spring 2026 three Claude Code changes made the same models worse for weeks (see “Compute and ‘getting dumber’”). More machinery isn’t automatically better: mini-SWE-agent, about 100 lines of Python with bash as its only tool, scores over 74% on SWE-bench Verified according to its authors. A leaderboard number describes a model–harness pair, so compare models in your own harness on your own tasks.
- The harness also sets most of the bill. Every turn resends the whole context, and in agents input tokens outnumber output roughly 100 to 1. So a harness orders each request from the most stable part to the most volatile: system prompt with tool definitions, then project memory, then the conversation, which only grows at the end. A cache read costs about 10% of the input price, and one change near the start of the prefix makes the whole history full price again (see “Prompt caching”). That is why Claude Code appends plan mode and skills as messages and applies a CLAUDE.md edit only after
/clear,/compactor a restart, while switching models or compacting rebuilds the cache.
Working with a coding agent
- Verification is the ground truth. Give the agent tests, a type checker, a linter and a build it can run itself, ideally a fast target such as
make test-fast. The agent’s “done” is a claim; the exit code is a fact. - Keep AGENTS.md small and precise: build, test and lint commands, conventions the code doesn’t reveal, and what not to touch. The Claude Code docs suggest under 200 lines, because a longer file costs context and is followed less reliably. Rules that must hold go into permissions and hooks, not prose.
- Plan before editing and scope the task: one bug or feature per session, with a clear criterion for done, and a fresh context between unrelated tasks.
- Review the diff, not the summary. Commit before a large change, so every agent edit can be reverted with one command.
- Match the permission level to the risk: read-only or plan mode in an unfamiliar repository, accepted edits inside a sandbox for daily work, and full autonomy (bypass, danger-full-access) only in a container or VM without secrets or production credentials.
Failure modes
- Losing the goal: after dozens of turns the task from the first message sits in the middle of the context or in a summary. A todo list, notes in files and shorter sessions help.
- Edits without tests: in LangChain’s traces the most common failure was an agent that wrote a solution, re-read it and stopped without running anything, and Anthropic (November 2025) saw agents mark features as finished without testing them end to end. A hook that runs the tests before the agent may finish helps.
- Context rot: quality drops as the history fills with logs and old file reads, long before the window is full (see “The context window and agents”). Compact at natural breaks between tasks and send searches to subagents.
- Prompt injection: a README, a code comment, an issue, a dependency’s docs or a web page is text in the same context as your instructions (see “Prompt injection”). In May 2025 Invariant Labs showed an issue in a public repository that made Claude 4 Opus, connected to the GitHub MCP server, read a private repository and publish its contents in a public pull request. The defence is in the harness: approval for actions, a network allowlist and tokens with minimal scope.
- Runaway cost: an agent retrying the same fix, subagents fanning out, dozens of tools loaded up front, a prefix that keeps changing. Set turn and budget limits in the harness and watch the ratio of cache reads to cache writes.
Check yourself
What does a harness add to the model in a coding agent, and why does the same model score differently in two harnesses?
A harness is everything around the model that makes it an agent: the loop, a few general tools (read, edit, shell, search), the system prompt and a project file such as AGENTS.md, a todo list, context management (clearing old results, compaction, subagents, tool search), skills loaded on demand, permissions and a sandbox, hooks and checkpoints. The model only proposes calls; the harness decides what it sees, what runs and what can be undone. So tools, prompts and a verification loop move scores: LangChain reports 13.7 more points on Terminal-Bench 2.0 after changing only the harness. The harness also keeps the prefix stable for prompt caching, which sets most of the bill.
Po polsku
Harness to wszystko wokół modelu, co robi z niego agenta: pętla, kilka ogólnych narzędzi (odczyt, edycja, powłoka, wyszukiwanie), system prompt i plik projektu, np. AGENTS.md, lista zadań, zarządzanie kontekstem (czyszczenie starych wyników, kompakcja, subagenci, wyszukiwanie narzędzi), skille ładowane na żądanie, uprawnienia i sandbox, hooki oraz checkpointy. Model tylko proponuje wywołania, a harness decyduje, co model widzi, co się wykona i co da się cofnąć. Dlatego narzędzia, prompty i pętla weryfikacji zmieniają wyniki: LangChain podaje 13,7 punktu więcej w Terminal-Bench 2.0 po zmianie samego harnessu. Harness trzyma też stały prefiks pod prompt caching, a od tego zależy większość rachunku.
Follow-up questions (5)
- Why doesn’t “never run git push” in AGENTS.md protect you?
- The file is context, not configuration. The model usually follows it, but a long session, an ambiguous request or injected text can outweigh it. A rule that must hold goes where the harness enforces it: a deny rule, a hook that blocks the call before it runs (in Claude Code this works even in bypass mode), a sandbox, or a token without push rights.
- The sandbox is on. How can the agent still destroy your work?
- The sandbox limits writes to the working directory and the network to allowed domains, and the repository is inside that boundary, so git clean -fdx or rm -rf src goes through. Claude Code checkpoints don’t track changes made by shell commands. What helps: frequent commits, an ask or deny rule for destructive commands, and a separate worktree or container for risky tasks.
- The agent keeps reporting success while CI fails. What do you change in the harness?
- A fast check the agent can run itself, with the command named in AGENTS.md. A hook for the moment the agent wants to finish that runs the tests and returns failures to the model as feedback. The result is judged by the exit code and the diff, not the summary. Anthropic and LangChain both describe a premature “done” as one of the most common failures of coding agents.
- Costs doubled after you connected three MCP servers. Why, and what do you do?
- Tool definitions sit in the prefix, so you pay for them on every turn, and in a harness that loads them up front, connecting a server mid-session invalidates the cache. What helps: tool search, so only names sit in the prefix, only the servers the task needs, and a fixed tool set for the whole session. Check the effect in usage: the ratio of cache reads to cache writes.
- When is a subagent better than doing the work in the main session?
- For side tasks that mostly read: searching the repository, reading logs, reviewing a diff. The subagent spends tokens in its own window and returns a short summary, so the main context stays short. Edits that depend on decisions from the main thread stay in the main session, because the subagent doesn’t see those decisions, and in Claude Code its edits usually don’t go into your checkpoints. In total, subagents cost more tokens, not fewer.