Home / Chapter 8 · Choosing a model
    Last edited · 7 min read

    Use with AI

    Compute and whether models “get dumber”

    The weights of a pinned model version do not change, but the system that serves them changes all the time: hardware, compilers, the sampler, routing, the app’s system prompt and its defaults. Behind an alias or a model name in a chat app, even the weights can be swapped while the name stays the same. On top of that, the provider has a finite pool of chips that both the training of the next model and today’s users draw from.

    In plain wordsOne kitchen with two orders at once: a big banquet next month, which is training, and today’s diners, who are the users. The number of burners stays the same. In a pinned version the recipe does not change, but a new oven or the wrong spice will change the taste.

    Move training and demand. See when inference runs out of capacity and what users see when it does

    trainingresearch and evalsinference in usefreerequests with no room

    Illustrative: 100 units of capacity, 10 of them permanently reserved for research and evals. The proportions are not any provider’s data.

    What changes while the weights stay put

    Documented cases

    Hypothesis and illusion

    How to protect yourself

    Check yourself

    Users say the model has “got dumber”. How would you explain that, and what would you do?

    First check whether the model changed at all: in an app or behind an alias the same name can point to new weights, as GPT-4o in ChatGPT did in April 2025. The weights of a pinned version do not change, but everything around them does: hardware and compilers, the sampler, routing, the app’s system prompt, default reasoning effort, rate limits. Both documented regressions at Anthropic, in 2025 and 2026, had such causes, not different weights. Silent quality cuts due to a compute shortage are technically possible but unproven. So instead of guessing, pin the version, set parameters explicitly, run your own evals on a schedule and compare traces.

    Po polsku

    Najpierw sprawdzasz, czy model w ogóle się zmienił: w aplikacji albo za aliasem ta sama nazwa może wskazywać nowe wagi, jak GPT-4o w ChatGPT w kwietniu 2025. Wagi przypiętej wersji się nie zmieniają, ale zmienia się wszystko wokół: sprzęt i kompilatory, sampler, routing, system prompt aplikacji, domyślny poziom rozumowania, limity. Oba udokumentowane spadki u Anthropic, z 2025 i 2026 roku, miały takie przyczyny, a nie inne wagi. Ciche cięcie jakości z braku mocy jest technicznie możliwe, ale niepotwierdzone. Dlatego zamiast zgadywać przypinasz wersję, ustawiasz parametry jawnie, cyklicznie puszczasz własne ewaluacje i porównujesz trace’y.

    Follow-up questions (4)
    How do you tell a regression at the provider from one of your own?
    A pinned version, a fixed eval set with several repetitions, and a comparison of traces. If the prompt, tools and parameters are the same and the scores have dropped beyond the noise, the problem is on the provider’s side. Then you collect examples and report them.
    What do you pin?
    The exact model snapshot ID, reasoning effort, the token limit, temperature where the model still accepts it, and the versions of the prompt and tool definitions. A pinned version also has a retirement date, so you run the migration against the same eval set.
    Why didn’t the provider catch the 2025 regression itself?
    According to the postmortem, Anthropic’s evals did not capture the drop, user reports were noisy, and privacy rules limit engineers’ access to conversations. The bugs hit part of the traffic on some platforms, so the averages looked fine. The lesson for you: you need evals on your own traffic.
    How do you detect a regression within an hour rather than a week?
    A canary: a small set of cases run every hour against the production endpoint, plus signals from traffic: the rate of invalid JSON, response length, number of tool calls, retries and user corrections. Alert on any deviation from the baseline.

    Sources

    Report an error · Suggest a fix