Home / Chapter 8 · Choosing a model
    Last edited · 10 min read

    Use with AI

    Choosing a model

    You choose a model for the task, not for the leaderboard: take the cheapest configuration, meaning model, reasoning effort and provider, that passes your eval set within your latency and limits. Public benchmarks only tell you which candidates are worth testing.

    In plain wordsHiring for a specific role. You don’t hire whoever scored highest on a general aptitude test; you give every candidate the same work sample from your actual job and take the one who does it well and on time for the lowest rate. Rankings and diplomas only tell you whom to invite for an interview. The work sample stays in the drawer, so when someone leaves you can vet a replacement in an hour.

    Set a quality floor and a latency cap. See which model is the cheapest of those that pass, then switch to the public leaderboard

    passescheapest that passesbelow the floortoo slowA, B: frontier · C, D, E: mid · F, G: small
    chosen model
    cost per completed task
    score on your eval set

    Illustrative data: seven unnamed models, numbers picked by hand. Cost per completed task: the cost of all attempts, failed ones and reasoning tokens included, after the cache discount, divided by the tasks passed. Under each dot: p95 latency of the whole task.

    From the task to the model

    Cost and latency

    Limits and data

    Reading benchmarks

    A choice that survives change

    Check yourself

    How do you choose a model for a new feature, and why doesn’t a public leaderboard settle it?

    Start from the task and your own eval set with a quality bar. Run candidates from several tiers, frontier, mid and small, on that set at different reasoning efforts, and take the cheapest configuration that clears the bar within the required latency. Count cost per completed task, including reasoning tokens, the model’s verbosity and the cache discount, rather than reading it off the per-token price list. Then check the limits: effective context length, output limit, rate limits, structured output, tool calling and data region. Public leaderboards can be contaminated, saturated, and sensitive to style and harness, so they only narrow the candidate list. Pin the version and keep the eval set in the repository, so that switching models is a single run.

    Po polsku

    Punktem wyjścia jest zadanie i własny zestaw ewaluacyjny z progiem jakości. Kandydatów z kilku półek, czołowej, średniej i małej, puszcza się na tym zestawie z różnymi poziomami rozumowania i bierze najtańszą konfigurację, która przechodzi próg w wymaganym czasie. Koszt liczy się na ukończone zadanie, z tokenami rozumowania, gadatliwością modelu i rabatem za cache, a nie z cennika za token. Do tego limity: efektywna długość kontekstu, limit wyjścia, limity zapytań, structured output, tool calling i region danych. Publiczne rankingi bywają skażone, nasycone, zależne od stylu i harnessu, więc tylko zawężają listę kandydatów. Wersję się przypina, a zestaw zostaje w repozytorium, żeby zmiana modelu była jednym uruchomieniem.

    Follow-up questions (5)
    The cheapest model scores 81% against an 80% bar. Is that enough?
    Not without checking the noise. With 50 cases the 95% confidence interval is about ±11 percentage points, so 81% and 80% are indistinguishable. You need more cases or repetitions, a paired comparison with a more expensive candidate, and a separate score for critical cases. If the cheap model fails on one type of case, that type can be routed to a stronger model.
    Model A costs five times more per token than model B. When does A come out cheaper?
    When it finishes the task in fewer steps: fewer agent turns, fewer fixes, shorter thinking at a lower effort, and more tasks passed. Cost per completed task is the cost of all attempts divided by the successful ones, so a model that passes 60% of tasks pays for 40% of its attempts for nothing. A measurement on your own set with full token accounting, cache and reasoning included, settles it.
    How do you build routing between a cheap and an expensive model?
    The simplest way is a cascade with a verifier: the cheap model answers, and a schema validator, tests or a judge decide whether to accept the result or rerun the task on a stronger model or at a higher reasoning effort. A router that predicts difficulty up front makes its own mistakes, so you measure its accuracy and the cost of its errors. In Anthropic’s measurements, running everything at low effort and rerunning failures at high effort held the score at about half the cost, but that only works where there is a reliable failure signal.
    A new model scores higher on a public benchmark. Should you switch?
    The public score only qualifies it as a candidate. What counts is your own eval set, the same one the current model passed, run pairwise with repetitions, plus cost per task, p95 latency and support for the features you use. The prompt usually needs retuning, so compare the best version for each model. Then a canary on part of the traffic, watching production signals.
    Why put an abstraction layer over the API if you use only one provider?
    Because providers retire versions, usually after a year or two, have outages and rate limits, and a better model may appear elsewhere. One layer keeps model names, prices, limits, fallback and token logging in one place, so a switch is configuration plus an eval run. It cannot hide feature differences such as tool formats, thinking or caching, so keep it thin.

    Sources

    Report an error · Suggest a fix