Home / Chapter 8 · Choosing a model
Last edited · 7 min read
Compute and whether models “get dumber”
The weights of a pinned model version do not change, but the system that serves them changes all the time: hardware, compilers, the sampler, routing, the app’s system prompt and its defaults. Behind an alias or a model name in a chat app, even the weights can be swapped while the name stays the same. On top of that, the provider has a finite pool of chips that both the training of the next model and today’s users draw from.
In plain wordsOne kitchen with two orders at once: a big banquet next month, which is training, and today’s diners, who are the users. The number of burners stays the same. In a pinned version the recipe does not change, but a new oven or the wrong spice will change the taste.
Move training and demand. See when inference runs out of capacity and what users see when it does
Illustrative: 100 units of capacity, 10 of them permanently reserved for research and evals. The proportions are not any provider’s data.
What changes while the weights stay put
- The same model runs on different hardware and software stacks. Anthropic says it serves Claude on AWS Trainium, NVIDIA GPUs and Google TPUs. A different precision, compiler or sampler implementation produces slightly different numbers at the output, and a bug in any of those places breaks only part of the traffic.
- A capacity shortage shows up first in latency and availability: queues, longer time to first token, slower tokens as batches grow, overload errors, stricter rate limits. It does not lower quality, but load sets the batch size, and with kernels that are not batch-invariant the same request can get different tokens even at temperature 0.
- Apps such as a chat or a coding tool change their system prompt, default reasoning effort and the way they manage context. That changes results without changing the model and without any change to the API.
Documented cases
- April 2025 (OpenAI, “Expanding on what we missed with sycophancy”). Here the weights did change, under the same name. On 25 April an update to GPT-4o in ChatGPT with new post-training, including an extra reward signal from thumbs-up and thumbs-down ratings, made the model noticeably sycophantic; the rollback began on 28 April. OpenAI counts five major post-training updates to GPT-4o in ChatGPT since its launch.
- August and September 2025 (Anthropic postmortem of 17 September 2025). Three overlapping infrastructure bugs, weights unchanged. From 5 August some Sonnet 4 requests went to servers configured for the upcoming 1M-token context, 16% of them at the worst hour on 31 August. From 25 August a misconfiguration of TPU servers caused out-of-place tokens to be inserted, e.g. Thai or Chinese characters in English answers. A change in token selection also exposed a bug in the XLA compiler for TPUs, through which approximate top-k sometimes dropped the most probable token. Fixes landed between 2 and 18 September.
- March and April 2026 (postmortem of 23 April 2026). Three changes in Claude Code, the Claude Agent SDK and Claude Cowork; the API was not affected. On 4 March the default reasoning effort was lowered from high to medium to cut latency; this was reverted on 7 April. On 26 March a bug in a mechanism meant to clear old thinking blocks once after an hour of inactivity made it clear them on every following turn, so the model seemed forgetful and repeated itself; fixed on 10 April. On 16 April a length limit on answers was added to the system prompt, which lowered an eval score by 3%; reverted on 20 April.
- According to both Anthropic postmortems, the causes were infrastructure bugs and product decisions made to reduce latency and response length, not a lack of capacity. In 2025 Anthropic stated plainly that it never reduces model quality due to demand, time of day or server load, and in 2026 that it never intentionally degrades its models.
Hypothesis and illusion
- The hypothesis: when chips run short before a launch, the provider quietly serves a more heavily quantised or smaller variant. That is technically possible at any provider, but there is no public evidence for it. Treat it as a hypothesis to test by measurement.
- The illusion: expectations rise, tasks get harder, sessions get longer, and a long context lowers quality on its own (see “The context window and agents”). With sampling, individual bad answers always happen, so anecdotes cannot tell a regression from noise.
How to protect yourself
- Pin an exact model snapshot ID, not an alias that can move; for Claude from 4.6 on, the dateless ID is itself the snapshot. Set reasoning effort and the token limit explicitly, and temperature only where the model still accepts it: newer Claude models (from Opus 4.7 on) and OpenAI models with reasoning enabled reject non-default values (see “The next token”). Version the prompt and the tool definitions (see “LLMs in production”).
- Run your own eval set on a schedule against the production endpoint, with several repetitions per case, because differences of a few percent get lost in the noise (see “Evals”). Production traces let you compare behaviour before and after.
- A pinned version also has a retirement date, so a migration will come anyway; run it against the same eval set. Full reproducibility needs an open model on a frozen stack of your own with deterministic, batch-invariant kernels (see “Open-weight models”).
Check yourself
Users say the model has “got dumber”. How would you explain that, and what would you do?
First check whether the model changed at all: in an app or behind an alias the same name can point to new weights, as GPT-4o in ChatGPT did in April 2025. The weights of a pinned version do not change, but everything around them does: hardware and compilers, the sampler, routing, the app’s system prompt, default reasoning effort, rate limits. Both documented regressions at Anthropic, in 2025 and 2026, had such causes, not different weights. Silent quality cuts due to a compute shortage are technically possible but unproven. So instead of guessing, pin the version, set parameters explicitly, run your own evals on a schedule and compare traces.
Po polsku
Najpierw sprawdzasz, czy model w ogóle się zmienił: w aplikacji albo za aliasem ta sama nazwa może wskazywać nowe wagi, jak GPT-4o w ChatGPT w kwietniu 2025. Wagi przypiętej wersji się nie zmieniają, ale zmienia się wszystko wokół: sprzęt i kompilatory, sampler, routing, system prompt aplikacji, domyślny poziom rozumowania, limity. Oba udokumentowane spadki u Anthropic, z 2025 i 2026 roku, miały takie przyczyny, a nie inne wagi. Ciche cięcie jakości z braku mocy jest technicznie możliwe, ale niepotwierdzone. Dlatego zamiast zgadywać przypinasz wersję, ustawiasz parametry jawnie, cyklicznie puszczasz własne ewaluacje i porównujesz trace’y.
Follow-up questions (4)
- How do you tell a regression at the provider from one of your own?
- A pinned version, a fixed eval set with several repetitions, and a comparison of traces. If the prompt, tools and parameters are the same and the scores have dropped beyond the noise, the problem is on the provider’s side. Then you collect examples and report them.
- What do you pin?
- The exact model snapshot ID, reasoning effort, the token limit, temperature where the model still accepts it, and the versions of the prompt and tool definitions. A pinned version also has a retirement date, so you run the migration against the same eval set.
- Why didn’t the provider catch the 2025 regression itself?
- According to the postmortem, Anthropic’s evals did not capture the drop, user reports were noisy, and privacy rules limit engineers’ access to conversations. The bugs hit part of the traffic on some platforms, so the averages looked fine. The lesson for you: you need evals on your own traffic.
- How do you detect a regression within an hour rather than a week?
- A canary: a small set of cases run every hour against the production endpoint, plus signals from traffic: the rate of invalid JSON, response length, number of tool calls, retries and user corrections. Alert on any deviation from the baseline.