## Choosing a model

*Choosing a model*

*Last edited: 28 September 2026*

You choose a model for the task, not for the leaderboard: take the cheapest configuration, meaning model, reasoning effort and provider, that passes your eval set within your latency and limits. Public benchmarks only tell you which candidates are worth testing.

**In plain words:** Hiring for a specific role. You don’t hire whoever scored highest on a general aptitude test; you give every candidate the same work sample from your actual job and take the one who does it well and on time for the lowest rate. Rankings and diplomas only tell you whom to invite for an interview. The work sample stays in the drawer, so when someone leaves you can vet a replacement in an hour.

*Interactive widget on the page: Set a quality floor and a latency cap. See which model is the cheapest of those that pass, then switch to the public leaderboard.*

### From the task to the model

- Start from a set of cases taken from real traffic and the bar it has to clear (see “Evals”). Run it on a few candidates, each case several times, and take the cheapest one that passes.
- Models come in tiers. Frontier models take hard reasoning, long agentic tasks and code. Mid-tier models cover most production work. Small, fast models handle classification, extraction, routing and simple steps at scale.
- Reasoning effort is the second knob of the same choice. In their model-selection guides Anthropic writes that tuning effort is often a better lever than switching models, and OpenAI advises keeping the lightest setting that meets your quality bar (as of September 2026). So you compare model and effort pairs, not bare models.
- One model for the whole system is rarely optimal. In a cascade a cheap model answers first, and only what a validator, tests or low confidence reject goes to a stronger model or a higher effort (see “LLMs in production”). Inside an agent you pick a model per step: a small one summarises tool results, a frontier one plans and makes the hard calls.

### Cost and latency

- Count cost per completed task, not price per token: all tokens from all attempts, failed ones included, divided by the number of tasks passed. A stronger model often finishes in fewer turns with fewer fixes. In Anthropic’s measurements on a SWE-bench Pro subset, Fable 5.1 at low effort solved 88.6% of tasks for 0.54 USD per solved task, while Sonnet 5, five times cheaper per token, solved 77.4% for 0.84 USD. On long research it went the other way: Fable 5.1 cost about four times more per task. These are vendor numbers, so check them on your own task.
- The price list hides three things. Reasoning tokens are billed as output, even when you never see them (see “Reasoning models”). Models differ in verbosity, so on the same task one can generate several times as many tokens as another. The cache discount differs too: at Anthropic a cache read usually costs 10% of the input price, 5% on Opus 5.5 and 2.5% on Fable 5.1 (as of September 2026, see “Prompt caching”). For an agent that reads most of its input from cache, that can flip the comparison.
- Latency is two numbers: time to first token (TTFT) and tokens per second after it. Reasoning effort moves both. The first word of the answer arrives only after all the thinking, and useful tokens per second drop because part of what is generated is thinking the user never sees. That is why Artificial Analysis measures time to first answer token separately. A slower model is helped by streaming, lower effort, or moving the step to where nobody is waiting.

### Limits and data

- The context window on the model card is an upper bound, not a promise: quality drops long before it is full (see “The context window and agents”), so test at your real lengths. The output limit is separate and much smaller, and thinking counts towards it.
- Rate limits depend on your account tier, which grows with your spending history. At Anthropic a new organisation may start on a lower tier, and each tier has a monthly spend cap after which the API returns 429 until the start of the next month (as of September 2026). Check the limits for the exact model before launch, not after.
- Strict structured output, parallel tool calls, images, logprobs, batch and prompt caching are not available on every model or at every provider (see “Enforcing output format”). One missing feature can rule out the model that wins on your eval.
- Where inference runs and where data is stored are two separate settings. Processing in a chosen region covers only some models and features, may need the provider’s approval, and costs more: about 10% on newer models at OpenAI and Anthropic (as of September 2026). Anthropic’s own API can pin only the US; an EU region comes through Bedrock or Google Cloud. Zero data retention (ZDR) needs an agreement and does not cover stateful features (see “LLMs in production”).

### Reading benchmarks

- Tasks from public test sets leak into training data, and once the top models approach the ceiling, differences vanish into noise and into errors in the set itself. In February 2026 OpenAI stopped reporting SWE-bench Verified: in the audited part of the set, 59.4% of tasks had flawed tests that rejected correct solutions, and every frontier model checked had seen some of the tasks and solutions in training.
- SWE-bench Verified is 500 bug fixes from 12 Python repositories. It measures small, well-described patches in popular projects, not work in your repository, with your conventions and changes across many files.
- LMArena measures which answer a voter prefers, and voters prefer longer, richly formatted ones. Once style was controlled for (style control, 2024), length turned out to be the strongest factor and the ranking reshuffled: GPT-4o-mini fell from 6th to 11th place, while Claude 3 Opus rose from 16th to 10th. It measures chat preference, not correctness on your task.
- A vendor-reported score and an independent one are different measurements. The vendor picks the harness, reasoning effort, number of attempts and task subset. Independent evaluators such as Artificial Analysis and Epoch AI run every model in one harness, and Epoch notes that SWE-bench results depend heavily on the scaffold. Compare only numbers from one source and one harness (see “The coding-agent harness”).
- The same model at different providers is not the same product: different quantisation, chat template, tool-call parser and serving settings. In Moonshot’s November 2025 measurement, Kimi K2 Thinking produced schema-valid tool calls 100% of the time through the official API and 83–87% of the time at several third-party hosts. Test the specific endpoint, not the model name.

### A choice that survives change

- Models are retired fast. Anthropic gives at least 60 days’ notice: Claude Sonnet 4, released in May 2025, got a retirement date in April 2026 and stopped working on 15 June 2026. Pin a specific version instead of an alias, and keep the eval set in your repository, so that switching models is one run plus a canary rather than a project (see “Compute and ‘getting dumber’”).
- Route model calls through one thin layer, your own or a gateway, that knows prices, limits and fallbacks. Then a change of model or provider is configuration. You will still have to retune the prompt, because a new model reads the same instructions differently.
- Open or closed is a separate decision about control over data, versions and cost at high volume (see “Open-weight models”). And before you pick a model, check whether the task needs an LLM at all (the next page, “When not to use an LLM: classifiers and System One”).

### Check yourself

**Question:** How do you choose a model for a new feature, and why doesn’t a public leaderboard settle it?

**Short answer:** Start from the task and your own eval set with a quality bar. Run candidates from several tiers, frontier, mid and small, on that set at different reasoning efforts, and take the cheapest configuration that clears the bar within the required latency. Count cost per completed task, including reasoning tokens, the model’s verbosity and the cache discount, rather than reading it off the per-token price list. Then check the limits: effective context length, output limit, rate limits, structured output, tool calling and data region. Public leaderboards can be contaminated, saturated, and sensitive to style and harness, so they only narrow the candidate list. Pin the version and keep the eval set in the repository, so that switching models is a single run.

### Follow-up questions

- **The cheapest model scores 81% against an 80% bar. Is that enough?** Not without checking the noise. With 50 cases the 95% confidence interval is about ±11 percentage points, so 81% and 80% are indistinguishable. You need more cases or repetitions, a paired comparison with a more expensive candidate, and a separate score for critical cases. If the cheap model fails on one type of case, that type can be routed to a stronger model.
- **Model A costs five times more per token than model B. When does A come out cheaper?** When it finishes the task in fewer steps: fewer agent turns, fewer fixes, shorter thinking at a lower effort, and more tasks passed. Cost per completed task is the cost of all attempts divided by the successful ones, so a model that passes 60% of tasks pays for 40% of its attempts for nothing. A measurement on your own set with full token accounting, cache and reasoning included, settles it.
- **How do you build routing between a cheap and an expensive model?** The simplest way is a cascade with a verifier: the cheap model answers, and a schema validator, tests or a judge decide whether to accept the result or rerun the task on a stronger model or at a higher reasoning effort. A router that predicts difficulty up front makes its own mistakes, so you measure its accuracy and the cost of its errors. In Anthropic’s measurements, running everything at low effort and rerunning failures at high effort held the score at about half the cost, but that only works where there is a reliable failure signal.
- **A new model scores higher on a public benchmark. Should you switch?** The public score only qualifies it as a candidate. What counts is your own eval set, the same one the current model passed, run pairwise with repetitions, plus cost per task, p95 latency and support for the features you use. The prompt usually needs retuning, so compare the best version for each model. Then a canary on part of the traffic, watching production signals.
- **Why put an abstraction layer over the API if you use only one provider?** Because providers retire versions, usually after a year or two, have outages and rate limits, and a better model may appear elsewhere. One layer keeps model names, prices, limits, fallback and token logging in one place, so a switch is configuration plus an eval run. It cannot hide feature differences such as tool formats, thinking or caching, so keep it thin.

### Sources

- [Anthropic: Optimizing for cost and intelligence (cost per completed task, effort, cascades)](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence)
- [LMSYS: Does style matter? Style control in Chatbot Arena (2024)](https://lmsys.org/blog/2024-08-28-style-control/)
- [Epoch AI: SWE-bench Verified benchmark review (with OpenAI’s 2026 audit)](https://epoch.ai/benchmarks/swe-bench-verified/review)
- [Artificial Analysis: how TTFT and output speed are measured](https://artificialanalysis.ai/methodology/performance-benchmarking)
- [Moonshot AI: K2 Vendor Verifier (one model, many hosts)](https://github.com/MoonshotAI/K2-Vendor-Verifier)

Interactive page: https://howaiworks.dev/choosing-a-model/
