## Tools (function calling)

*Agents*

*Last edited: 28 September 2026*

The model runs no code. It gets a list of tools, each with a name, a description and an argument schema, and when it decides it needs one, it writes a call instead of an answer: your code runs it and sends back the result.

**In plain words:** A manager who can only make phone calls. They have a card listing the jobs an assistant can handle, with a short description of each. They say “check whether tomorrow’s 8:15 train is running on time”, and the assistant checks and calls back with the result. When the descriptions on the card are vague, the manager asks for the wrong thing.

*Interactive widget on the page: Change the quality of the tool descriptions and toggle parallel calls, then step through the exchange. Watch which tool the model reaches for, how many requests it costs and whether the answer is true.*

### What a call looks like

- The model has no access to the internet, a database or your computer. All it can do is write, in an agreed format, “call weather_forecast for Manchester for tomorrow”. The result comes back to it as another message, and it reads it like any other text. How this turns into multi-step work is covered in “The agent loop”.
- Definitions go into the prompt as text in a format the model knows from training, like the roles in “How a model sees a chat”. A call is ordinary tokens too: the API extracts it and returns it as a separate field with an ID, a name and arguments, and the stop reason says the model is waiting for a result. You run several calls from one response concurrently and send back all the results, each with its call’s ID. When order matters, you turn parallel calls off (`parallel_tool_calls: false` in OpenAI, `disable_parallel_tool_use` in Anthropic).
- `tool_choice`: `auto` (the default, the model decides), `required` or `any` (it must call one of them), a specific tool (e.g. for extracting data into a schema), `none` (calls not allowed). Forcing a call on every turn won’t let the model finish, so in a loop it is used only at selected steps. Forcing doesn’t work everywhere (as of September 2026): Claude Opus 5.5 and Fable 5.1 reject `any` and a specific tool with a 400 error. To extract data into a schema there, use structured outputs (see “Enforcing output format”).
- Some tools are run by the provider: web search or running code in a sandbox happens on their side, and you get the finished result. Less code to write, less control over what ran and where.
- Computer use and browser agents run the same loop with a screen as the tool: the model gets a screenshot, writes an action (click at x, y, type text), your code performs it and returns a fresh screenshot. Each screenshot is image input, about 1,000–1,800 tokens at Anthropic, and stays in the history, so long sessions prune old ones. Every page the agent looks at is untrusted input: text on screen can steer it like any tool result (see “Prompt injection”).

### The description is the only manual

- The model chooses a tool only by its name, description and schema; it never sees the code. Write the description as you would for a new team member: what the tool does, when to use it, when not to, what format the arguments take and what it returns. For the model, two tools with similar descriptions are a coin toss, as the overlapping-tools variant shows.
- What helps: names prefixed with the service (`rail_status`, `rail_timetable`), unambiguous parameters (`user_id` rather than `user`), enums instead of free text, and example calls: in Anthropic’s tests, examples in the definition raised accuracy on complex arguments from 72% to 90%. A few tools built for specific tasks beat a wrapper around every API endpoint.
- You design the result too. Return only the fields that are needed, readable names instead of internal UUIDs, and trim long lists with a note on how to fetch the rest. Every token of a result stays in the context and is sent on every later turn, unless your harness clears or compacts it (see “Context engineering and memory”).
- Strict mode (`strict`) is the constrained decoding from “Enforcing output format”: the arguments always match the schema. It doesn’t guarantee sensible values: the model can still guess a missing parameter or pick the wrong day.

### Every tool costs on every request

- Definitions are text attached to every request, so you pay for them on every turn, even when no tool gets used. The provider adds its own hidden system prompt for tool use on top: a few hundred tokens at Anthropic, depending on the model.
- Definitions sit at the start of the prompt, so they cache well (see “Prompt caching”). Adding, removing or reordering a tool mid-conversation invalidates the cache for everything after them. OpenAI has `allowed_tools` for this: it narrows the choice on a given turn without changing the list of definitions.
- More tools mean a higher bill, less room in the context and a harder choice. OpenAI recommends fewer than 20 functions at a time, noting that this is a soft suggestion. Anthropic gives an example of five MCP servers with 58 tools whose definitions took about 55k tokens before the first question was asked (see “MCP”). At that scale tool search helps: the model sees a tool for searching tools and gets the full definitions only when it needs them.

### Errors and security

- An error is a result too. Instead of aborting, send the model a specific message: what is wrong and how to fix it, and the model usually corrects itself on the next step. Check the arguments in code before running anything. Protect state-changing operations with an idempotency key: after a network error, your code or the model may retry the same call, and the ticket must not be bought twice.
- A tool is a permission. The model can call it with wrong arguments or under the influence of text it has just read, because tool results (web pages, emails, documents) are untrusted data. Grant the least privilege needed, and make irreversible actions, such as a transfer, a deletion or a send, wait for human confirmation, enforced in code, not in the prompt. More in “Prompt injection”.

### Check yourself

**Question:** How does function calling work, and how do you design tools for a model?

**Short answer:** The model executes nothing. With the prompt it gets tool definitions: a name, a description and a JSON Schema for the arguments. When it needs a tool, it emits a call in a trained format instead of an answer, sometimes several at once. Your code validates the arguments, runs the call and returns the result with the call ID. Selection depends mostly on the descriptions, so write them like documentation, avoid overlapping tools and keep results concise. Definitions cost tokens on every request. Errors go back as results with a hint, state-changing operations are idempotent, and irreversible ones need human approval.

### Follow-up questions

- **How does function calling differ from structured output?** From the model’s side it is the same mechanism: text in a prescribed format. The difference is in what happens next. Structured output is the final answer for your code. A tool call is a request, after which the result goes back to the model and the conversation continues.
- **How many tools is too many?** There is no hard limit. The signals are wrong tool choices on your eval set and the growing cost of definitions. Then you merge similar tools, split the work between sub-agents with their own tool sets, or load definitions on demand through tool search.
- **What do you do when the model passes wrong arguments?** First, improve the description and the schema: format, examples, enums. Then validate in code and send a readable error back to the model so it can correct itself. Strict mode removes syntax and type errors, but not wrong values.
- **A tool result is 50k tokens. What do you do?** You don’t paste it in whole, because it would stay in the context and be sent on every turn until something clears it. You filter the fields, paginate, return a summary with a handle to the full data, and write large content to a file the agent can search with a separate tool.
- **When do you give the model code execution instead of many calls?** When the task means many calls plus data processing, e.g. fetch 50 records, filter them, count them. The model writes a script that calls the tools itself, and only the final result comes back into the context. Anthropic reports about 37% fewer tokens on complex research tasks. The price is a sandbox for running code and harder auditing.

### Sources

- [Anthropic: Writing effective tools for agents](https://www.anthropic.com/engineering/writing-tools-for-agents)
- [OpenAI: Function calling (docs)](https://developers.openai.com/api/docs/guides/function-calling)
- [Anthropic: Advanced tool use (tool search)](https://www.anthropic.com/engineering/advanced-tool-use)
- [Claude docs: Define tools (tool_choice and limits on forcing)](https://platform.claude.com/docs/en/agents-and-tools/tool-use/define-tools)
- [Claude docs: Computer use tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool)

Interactive page: https://howaiworks.dev/tool-calling/
