Home / Chapter 5 · Agents
    Last edited · 8 min read

    Use with AI

    Multiple agents

    A lead agent splits the task and delegates parts to subagents, each working in its own clean context and returning a short result. It helps with independent parts, such as researching many sources at once, and hurts with shared state, such as a change to one codebase. It usually costs several times more tokens.

    In plain wordsFive researchers will get through ten libraries faster than one, as long as each gets a clear brief and hands back a page of notes. Five programmers fixing the same module without talking to each other will do more damage than one, because each will quietly make different decisions. You pay both teams for every hour.

    Change the number of subagents and the type of task. Watch the time, the tokens, the lead agent’s context and the quality of the result

    from request to result
    largest context of the lead agent

    Illustrative numbers, picked by hand, not measured. Tokens are the input and output of all agents combined. The number to the right of each bar is that agent’s largest context; a gap in the bar is waiting. Assumptions: in the research task one library takes an agent 4 minutes, in the code task the dependencies allow at most two parallel paths, and every pair of agents diverges on one shared decision.

    Lead agent and subagents

    When it helps

    When it hurts

    How to build it

    Check yourself

    When would you build a system with multiple agents, and when would you stick with one?

    Start with a single agent. Add more when the task splits into independent parts, such as researching many sources: a lead agent hands them to subagents, and each works in parallel in its own clean context and returns a short summary. The gain is speed, breadth and a clean lead context. The price is tokens, about 15 times a chat in Anthropic’s data, and coordination. With shared state, such as one change across a codebase, agents make conflicting decisions, so there one agent makes the changes. End-state evals decide.

    Po polsku

    Zaczyna się od jednego agenta. Kolejnych dodaje się, gdy zadanie rozpada się na niezależne części, jak badanie wielu źródeł: główny agent rozdziela je subagentom, każdy pracuje równolegle we własnym, czystym kontekście i oddaje krótkie streszczenie. Zysk to czas, szerokość i czysty kontekst głównego agenta. Cena to tokeny, u Anthropic ok. 15 razy więcej niż czat, oraz koordynacja. Przy wspólnym stanie, jak zmiana w jednym kodzie, agenci podejmują sprzeczne decyzje, więc tam zmiany robi jeden agent. Rozstrzygają ewaluacje stanu końcowego.

    Follow-up questions (5)
    A subagent doesn’t know what the others decided. How do you deal with that?
    Either pass it the full context and the decisions made so far, which eats up the gain from isolation, or agree the shared things before the split: names, types, output format, the scope of each part. If that can’t be settled up front, the parts aren’t independent and one agent will do better.
    The lead agent spawns 50 subagents for a simple question. What do you do?
    That was a real bug in an early version of Anthropic’s system. Scaling rules go into the prompt: how many subagents and calls for which type of question. Code adds hard limits on the number of subagents, the depth and the token budget, and evals check the change.
    How do you pass large results between agents?
    Through files or a store: the subagent writes an artefact and returns its path with a short description. Details don’t get lost in successive summaries, and the lead agent’s context stays small.
    How is a handoff different from calling a subagent?
    A subagent works like a tool: the lead agent waits for the result and keeps control. A handoff passes the conversation, with its state, to another agent, which carries on talking to the user itself. A handoff suits routing a case to a specialist; a subagent suits splitting a task into parts.
    How do you debug and evaluate a multi-agent system?
    One trace ID for the whole task and a separate span for each subagent with its brief, calls and tokens. Grade the end state on a fixed set and compare it with a single agent, because a small change to the lead agent’s prompt can change the behaviour of every subagent.

    Sources

    Report an error · Suggest a fix