Home / Chapter 5 · Agents
    Last edited · 8 min read

    Use with AI

    Tools (function calling)

    The model runs no code. It gets a list of tools, each with a name, a description and an argument schema, and when it decides it needs one, it writes a call instead of an answer: your code runs it and sends back the result.

    In plain wordsA manager who can only make phone calls. They have a card listing the jobs an assistant can handle, with a short description of each. They say “check whether tomorrow’s 8:15 train is running on time”, and the assistant checks and calls back with the result. When the descriptions on the card are vague, the manager asks for the wrong thing.

    Change the quality of the tool descriptions and toggle parallel calls, then step through the exchange. Watch which tool the model reaches for, how many requests it costs and whether the answer is true

    User: What will the weather be like in Manchester tomorrow morning, and will the 8:15 train to London run on time?

    The definitions sent along with the question: name, arguments and description.

    The bars show how the model spreads its odds across the tools for each part of the question. In this run it picked the tallest bar; with a different sample it could pick another one (see “The next token”). Illustrative numbers, not a measurement of a real model.

    definition tokens in every request
    requests to the model
    tool calls
    tokens spent on definitions alone in this conversation

    A simplified, provider-neutral format: no API accepts it as is, and each names the fields differently (OpenAI, for example, returns arguments as a JSON string). Tokens are estimated from the length of the minified definition JSON: about 4 characters per token for English descriptions, closer to 3 for Polish.

    What a call looks like

    The description is the only manual

    Every tool costs on every request

    Errors and security

    Check yourself

    How does function calling work, and how do you design tools for a model?

    The model executes nothing. With the prompt it gets tool definitions: a name, a description and a JSON Schema for the arguments. When it needs a tool, it emits a call in a trained format instead of an answer, sometimes several at once. Your code validates the arguments, runs the call and returns the result with the call ID. Selection depends mostly on the descriptions, so write them like documentation, avoid overlapping tools and keep results concise. Definitions cost tokens on every request. Errors go back as results with a hint, state-changing operations are idempotent, and irreversible ones need human approval.

    Po polsku

    Model niczego nie wykonuje. Z pytaniem dostaje definicje narzędzi: nazwę, opis i schemat argumentów w JSON Schema. Gdy uzna narzędzie za potrzebne, zamiast odpowiedzi wypisuje wywołanie w wyuczonym formacie, czasem kilka naraz. Twój kod waliduje argumenty, wykonuje wywołanie i odsyła wynik z id wywołania. Model wybiera głównie po opisie, więc opisy pisze się jak dokumentację, bez nakładających się narzędzi, a wyniki są zwięzłe. Za definicje płaci się w każdym requeście. Błąd wraca jako wynik ze wskazówką, operacje zmieniające stan są idempotentne, a nieodwracalne zatwierdza człowiek.

    Follow-up questions (5)
    How does function calling differ from structured output?
    From the model’s side it is the same mechanism: text in a prescribed format. The difference is in what happens next. Structured output is the final answer for your code. A tool call is a request, after which the result goes back to the model and the conversation continues.
    How many tools is too many?
    There is no hard limit. The signals are wrong tool choices on your eval set and the growing cost of definitions. Then you merge similar tools, split the work between sub-agents with their own tool sets, or load definitions on demand through tool search.
    What do you do when the model passes wrong arguments?
    First, improve the description and the schema: format, examples, enums. Then validate in code and send a readable error back to the model so it can correct itself. Strict mode removes syntax and type errors, but not wrong values.
    A tool result is 50k tokens. What do you do?
    You don’t paste it in whole, because it would stay in the context and be sent on every turn until something clears it. You filter the fields, paginate, return a summary with a handle to the full data, and write large content to a file the agent can search with a separate tool.
    When do you give the model code execution instead of many calls?
    When the task means many calls plus data processing, e.g. fetch 50 records, filter them, count them. The model writes a script that calls the tools itself, and only the final result comes back into the context. Anthropic reports about 37% fewer tokens on complex research tasks. The price is a sandbox for running code and harder auditing.

    Sources

    Report an error · Suggest a fix