## LLMs in production

*Production and serving*

*Last edited: 28 September 2026*

A model call is a remote dependency that is sometimes overloaded, slow and expensive, and sometimes returns something other than what you asked for. A production system is the code around that call: retries and fallbacks, control of latency and cost, traces, validation and safe rollout of changes.

**In plain words:** A restaurant with one brilliant but temperamental chef: sometimes swamped with orders, sometimes slow, sometimes sending out the wrong dish. The floor manager can’t fix the chef, but can put the order in again a moment later, keep a stand-in chef ready, check the plate before it goes out and note what went wrong.

*Interactive widget on the page: Pick a failure and turn on safeguards. See what the user gets, how long it takes and what it costs.*

### Failures: what to retry and how

- Retry only transient errors: 429 (rate limit), overload (529 at Anthropic, 503 at OpenAI), other 5xx, dropped connections and timeouts. A 400, 401 or 403 will fail the same way the second time. So will a 429 from your tier’s monthly spend cap: at Anthropic it has no `retry-after` header, and `error.details.error_code` is `enforced_spend_limit_reached`.
- Back off exponentially with random jitter: about 1, 2, 4 s plus a random extra, and when the server sends `Retry-After`, at least that long. Without jitter all clients come back in the same second and knock the overloaded service down again. Cap the number of attempts and the total time. Retry in one layer only: the official SDKs retry on their own (Anthropic’s twice by default), so your own three-attempt loop on top of the SDK turns one click into up to 9 requests.
- A retry is safe only when repeating the call breaks nothing. The model call itself changes nothing, but an agent’s tool does. Every action with a side effect gets an idempotency key (for example the task ID plus the step number), and the tool rejects duplicates, so a retried step won’t send a second email. Don’t automatically repeat a response you have already started streaming to the user.
- Rate limits are counted in requests and tokens per minute (RPM and TPM; at Anthropic, input and output separately). Anthropic replenishes them continuously with a token bucket, and a 60 RPM limit can behave like 1 request per second, so a sudden burst of traffic gets 429s despite headroom on the per-minute scale. OpenAI counts `max_tokens` towards the limit if it exceeds the request’s estimate. Anthropic counts the tokens actually generated and, for most models, does not count cache reads at all. The limit is shared by the whole organisation, so you need your own queue and per-user limits to stop one customer from blocking everyone else.
- The closest fallback is the same model on a platform someone else runs: Claude on Bedrock or Google Cloud, GPT on Azure or Bedrock (as of September 2026). It is a separate account with its own quotas and model IDs, some features arrive there later, and the cache is cold. A different model also saves availability, but a prompt tuned for the primary behaves differently on it and the tool format may differ. Either way the fallback path needs its own evals, because a rarely used path breaks silently. A circuit breaker, after a run of errors, briefly stops calling the failing provider and sends traffic straight to the fallback, letting a trial request through every so often. When nothing works, degrade gracefully: a cached result, a simpler answer or a clear message instead of a spinner that never stops.

### Latency

- Two numbers matter: time to first token (TTFT, which is queueing plus prompt processing) and total time (TTFT plus the number of output tokens times the time per token). Streaming shortens only the perceived wait: the user reads from the first token, but the whole answer takes just as long. With streaming, set the timeout on silence between chunks, not on the whole response, and treat an error event mid-stream (it can arrive after 200 OK) as a failure, not as the end of the answer. A long response without streaming risks the idle connection being dropped. The price of streaming is validation: you can check JSON only once it is complete.
- Latency grows mainly with output length, because every token is a separate decoding step (see “Why the GPU is idle”). By OpenAI’s rule of thumb, halving the output roughly halves latency, while halving the prompt cuts it by only 1–5%. Reasoning tokens are output too. A concise format, short field names and a lower reasoning effort where it isn’t needed all help.
- Run independent calls in parallel, for example question classification and retrieval. Give simple steps such as routing, classification or extraction to a smaller, faster model. Whatever a rule or plain code can handle, do without a model.

### Cost

- Put the fixed prefix (system prompt, tools, documents) first, because a cache read costs a fraction of the input price, usually 10% at Anthropic (see “Prompt caching”). Run offline jobs such as nightly classification, data enrichment or eval runs through the batch API: 50% cheaper at OpenAI and Anthropic, with results within 24 h (as of September 2026).
- Cascade: a cheap model first, and the expensive one only when a validator rejects the result or confidence is low. It pays off when most traffic is simple and you can cheaply check whether the cheap model coped. Without such a check it is just a quality downgrade.
- Hard limits: `max_tokens` on every call (it is a time limit too), a maximum number of agent steps, a daily budget per user or customer, and a cost alert. Store the results of repeatable tasks in an ordinary cache keyed on the input, prompt version and model version. A semantic cache, which matches similar questions, can return the answer to a different question.

### Observability

- Every task is a trace, and every model call and tool step is a span within it: input, output, tokens (including cached ones), TTFT and total time, cost, the exact model version from the response, the prompt version and the provider’s request ID. OpenTelemetry has GenAI conventions for this (`gen_ai.*` attributes, still in development).
- Metrics: p50 and p95 latency (the mean hides the tail), error rate by type (429, 5xx, timeout, validation), fallback share, cost per completed task rather than per call, and online eval results: an LLM judge on a sample of traffic and user ratings.
- Prompts and responses contain personal data and company secrets. Mask PII before storing them, and restrict access to logs and how long they are kept. In the OpenTelemetry conventions, capturing message content is off by default. Turn traces where the system failed into test cases (see “Evals”), so every incident becomes a regression test.

### Guardrails and rolling out changes

- Check output in code: schema, types and business rules (the amount is not larger than the balance, the ID exists). On failure, send the model a specific error message and retry once or twice, then fall back or return an error. Strict mode removes syntax errors, not bad values (see “Enforcing output format”). Where the domain requires it, add an input length limit, moderation and PII masking. Risky actions such as payments, deletions or sending anything externally are approved by a human, and code enforces that, not the prompt.
- Version prompts like code and pin the exact model version, not an alias that can move. Every change to the prompt, model or parameters goes through offline evals first, then a canary on a few percent of traffic with a metric comparison and a quick rollback. Providers retire old versions, so a migration will come anyway, and your own evals are the only proof that the new version is not worse (see “Compute and ‘getting dumber’”).
- Know where the data goes. At OpenAI and Anthropic, content sent through the API is retained by default for up to 30 days (abuse monitoring), and for less only under a zero data retention (ZDR) agreement. At Anthropic, Claude Fable and Mythos 5.x require 30-day retention and are excluded from ZDR, so the model choice is also a data decision. Stateful features such as files, batch results or stored responses keep it longer (as of September 2026). Your logs and tracing tool are a second copy of the same data.
- EU AI Act (as of September 2026): since 2 August 2026 a chatbot must tell people they are talking to an AI, and generated content must carry a machine-readable mark (systems already on the market have until 2 December 2026). High-risk uses such as CV screening, credit scoring or grading exams get their obligations from 2 December 2027.

### What the user sees

- Show progress and let the user stop it: during reasoning or agent steps, show which step is running (searching, reading a file) next to a stop button. A minute-long spinner looks like a hang, and a wrong path is cheapest to stop early.
- Show where the answer comes from: sources linked to the passages they support, and a plain “nothing found” instead of an answer from memory. Signal uncertainty with what you can check (no sources, a failed validator, samples that disagree), not with the confidence the model states about itself (see “Hallucinations”).
- Agent actions get a preview before and an undo after: a draft instead of a sent email, soft delete, a change log with a revert. Only what cannot be undone needs the human approval described above.

### Check yourself

**Question:** You are shipping an LLM-based feature to production. What do you build around the model call itself?

**Short answer:** Treat the model as an unreliable, slow and expensive external dependency. Every call has a timeout, and transient errors (429, overload, 5xx) are retried with exponential backoff and jitter, in one layer only. Tool actions carry idempotency keys. The fallback, ideally the same model on another platform, gets its own evals and a circuit breaker. Streaming, shorter outputs, prompt caching, a cheaper model for simple steps and offline batch jobs cut latency and cost. Every call lands in a trace with tokens, cost and version. Outputs are validated in code, and prompt and model changes ship through evals and a canary.

### Follow-up questions

- **The provider has been returning 529 for ten minutes. What does the user see?** After a run of errors the circuit breaker stops calling the provider and traffic goes straight to the fallback, so the user gets the backup model’s answer without waiting through more retries. Without a fallback: a fast, clear message, and tasks that can wait go into a queue. When trial requests succeed, the breaker gradually restores traffic.
- **Why not retry every error?** A 400 or 401 will fail the same way the second time, and a 429 from an exhausted spend limit won’t clear on its own. Blind retries multiply traffic at the worst possible moment: under overload, every client making three attempts triples the load. Hence jitter, a limit on total time and retries in one layer only.
- **How would you halve the bill without losing quality?** Measure first: cost per completed task, broken down into input, output and cache. Then a stable prefix for prompt caching, batch for everything non-interactive, shorter outputs, a cheaper model where evals show no difference, and a cache for the results of repeatable tasks. Check every change against evals, because a cheaper model can quietly lower quality.
- **How do you reconcile streaming with JSON validation?** Full validation is possible only after the last token. In the UI, stream text or parse the structure incrementally and show the fields that are already complete. Where an error is costly, for example before a write or an action, buffer and validate the whole thing.
- **How do you move safely to a newer model?** Offline evals on a set built from real cases, a comparison of cost and latency, then a canary on a few percent of traffic with the same metrics and a rollback ready. The prompt usually needs retuning, because a new model reads the same instructions differently. Test a fallback to a different model the same way, because that is a model change too.

### Sources

- [OpenAI: Rate limits (retrying with exponential backoff and jitter)](https://developers.openai.com/api/docs/guides/rate-limits)
- [Claude docs: Errors (which errors to retry)](https://platform.claude.com/docs/en/api/errors)
- [OpenAI: Latency optimization](https://developers.openai.com/api/docs/guides/latency-optimization)
- [OpenTelemetry: GenAI semantic conventions](https://github.com/open-telemetry/semantic-conventions-genai)
- [EUR-Lex: Digital Omnibus on AI, Regulation (EU) 2026/1744 (AI Act dates)](https://eur-lex.europa.eu/eli/reg/2026/1744/oj/eng)

Interactive page: https://howaiworks.dev/llms-in-production/
