Home / Chapter 5 · Agents
    Last edited · 9 min read

    Use with AI

    The agent loop

    An agent is a model in a loop: it gets a goal and tools, picks a call, your code runs it and sends back the result, until the model answers without a call. It differs from a workflow in one thing: the model, not your code, chooses the next step.

    In plain wordsA workflow is a checklist for an intern: do this, then that, and if something doesn’t fit, come to me. An agent is an assistant who gets a goal, a phone and a calendar, and works out the next steps alone. It copes with surprises, but you don’t know in advance how long it will take or what it will do along the way.

    Choose workflow or agent and what is in the calendar on Friday, then press ▶. Watch who picks the next step and how the context grows with every turn

    model calls
    tool calls
    input tokens in total

    The bar shows what the last model call is made of (scale: 3,500 tokens). Token counts are illustrative, not measured. The calendar, tools and model responses are made up to show the flow.

    The loop in code

    Workflow or agent

    Failures in long tasks

    Cost, permissions and evaluation

    Check yourself

    When would you build an agent instead of a workflow, and how do you secure its loop in production?

    An agent is a model in a loop: it gets a goal and tools, picks a call, your code executes it and appends the call and its result, until the model answers without a call. In a workflow you write the steps and the model does narrow jobs inside them, so it is cheaper, faster and repeatable. An agent pays off only when the steps cannot be listed up front. Then code enforces turn, budget and time limits, grants least privilege, holds irreversible actions for human approval and checkpoints state after every step. Quality is measured by the end state, with each task run several times.

    Po polsku

    Agent to model w pętli: dostaje cel i narzędzia, wybiera wywołanie, twój kod je wykonuje i dopisuje wywołanie z wynikiem, aż model odpowie bez wywołania. W workflow kroki zapisujesz ty, a model robi w nich wąskie zadania, więc jest taniej, szybciej i powtarzalnie. Agent opłaca się dopiero, gdy kroków nie da się wypisać z góry. Wtedy kod pilnuje limitu tur, budżetu i czasu, daje najmniejsze uprawnienia, wstrzymuje akcje nieodwracalne do zgody człowieka i zapisuje stan po każdym kroku. Jakość mierzy się stanem końcowym, a każde zadanie puszcza kilka razy.

    Follow-up questions (6)
    How does an agent know it is done?
    It doesn’t know, it judges: it answers without calling a tool. Sometimes it announces a success that never happened. That is why code has hard limits (turns, tokens, time), and you check the result with an end-state test, not the model’s summary.
    The agent keeps calling the same tool. What do you do?
    As a quick fix, code detects a repeated call with the same arguments, adds a note for the model, and breaks the loop on the next repeat. The cause is usually in the tool: an unclear error message or a result that shows no progress. The run goes into the eval set as a new case.
    The agent stopped halfway through a task. How do you design for that?
    You assume it will happen. State-changing tools accept an idempotency key, and you save the loop state after every call, so work can resume without repeating side effects. A human gets a list of what has already been changed and, where possible, an action that undoes it.
    How do you judge whether an agent is ready for production?
    A set of tasks from real cases, each run several times and graded on the end state, plus cost and time per task. You compute pass^k, because users expect it to work every time: for a task the agent completes in 90% of attempts, with independent attempts, three successes in a row happen in about 73% of cases. Across a suite you compute pass^k per task and average it.
    How do you classify an agent’s failures?
    Following Chip Huyen (2025), into three groups. Planning: a wrong or non-existent tool, bad arguments, a missed goal, a false claim of success. Tool: the call is right, but the tool returned a wrong result, so each tool is tested on its own. Efficiency: the task is done, but with too many steps, too much cost or too much time against a baseline. Traces tagged this way show whether to fix the prompt and tool descriptions, the tool itself or the limits.
    Do you need a framework to build an agent?
    No. The loop is a dozen or so lines on a plain API, and that is the place to start. A framework helps with durable state, resuming and traces, but it hides the prompt and the context, which you have to understand anyway.

    Sources

    Report an error · Suggest a fix