Home / Chapter 7 · Production and serving
    Last edited · 11 min read

    Use with AI

    LLMs in production

    A model call is a remote dependency that is sometimes overloaded, slow and expensive, and sometimes returns something other than what you asked for. A production system is the code around that call: retries and fallbacks, control of latency and cost, traces, validation and safe rollout of changes.

    In plain wordsA restaurant with one brilliant but temperamental chef: sometimes swamped with orders, sometimes slow, sometimes sending out the wrong dish. The floor manager can’t fix the chef, but can put the order in again a moment later, keep a stand-in chef ready, check the plate before it goes out and note what went wrong.

    Pick a failure and turn on safeguards. See what the user gets, how long it takes and what it costs

    Failure

    Safeguards

      first content for the user
      request handled
      model calls
      cost, where 1× is one normal call

      An illustrative simulation with hand-picked numbers. With no failure: first token after 0.8 s, the full answer after 6.8 s, cost 1×. The fallback model at another provider has a cold cache, so it costs 1.4×. Policy: a 20 s timeout per attempt (with streaming, 10 s of silence between chunks), at most 2 retries within an overall limit of 45 s, fallback once retries are exhausted. An interrupted attempt costs in proportion to what it managed to generate.

      Failures: what to retry and how

      Latency

      Cost

      Observability

      Guardrails and rolling out changes

      What the user sees

      Check yourself

      You are shipping an LLM-based feature to production. What do you build around the model call itself?

      Treat the model as an unreliable, slow and expensive external dependency. Every call has a timeout, and transient errors (429, overload, 5xx) are retried with exponential backoff and jitter, in one layer only. Tool actions carry idempotency keys. The fallback, ideally the same model on another platform, gets its own evals and a circuit breaker. Streaming, shorter outputs, prompt caching, a cheaper model for simple steps and offline batch jobs cut latency and cost. Every call lands in a trace with tokens, cost and version. Outputs are validated in code, and prompt and model changes ship through evals and a canary.

      Po polsku

      Model traktuje się jak zawodną, wolną i drogą zależność zewnętrzną. Każde wywołanie ma timeout, a błędy przejściowe, czyli 429, przeciążenie i 5xx, ponawia się z wykładniczym backoffem i jitterem, w jednej warstwie. Akcje narzędzi mają klucze idempotencji. Fallback, najlepiej ten sam model na innej platformie, ma własne evale i circuit breaker. Czas i koszt obniżają streaming, krótsze wyjście, prompt caching, tańszy model do prostych kroków i batch offline. Każde wywołanie trafia do trace’a z tokenami, kosztem i wersją. Wyjście waliduje się w kodzie, a zmiany promptu i modelu wypuszcza przez evale i canary.

      Follow-up questions (5)
      The provider has been returning 529 for ten minutes. What does the user see?
      After a run of errors the circuit breaker stops calling the provider and traffic goes straight to the fallback, so the user gets the backup model’s answer without waiting through more retries. Without a fallback: a fast, clear message, and tasks that can wait go into a queue. When trial requests succeed, the breaker gradually restores traffic.
      Why not retry every error?
      A 400 or 401 will fail the same way the second time, and a 429 from an exhausted spend limit won’t clear on its own. Blind retries multiply traffic at the worst possible moment: under overload, every client making three attempts triples the load. Hence jitter, a limit on total time and retries in one layer only.
      How would you halve the bill without losing quality?
      Measure first: cost per completed task, broken down into input, output and cache. Then a stable prefix for prompt caching, batch for everything non-interactive, shorter outputs, a cheaper model where evals show no difference, and a cache for the results of repeatable tasks. Check every change against evals, because a cheaper model can quietly lower quality.
      How do you reconcile streaming with JSON validation?
      Full validation is possible only after the last token. In the UI, stream text or parse the structure incrementally and show the fields that are already complete. Where an error is costly, for example before a write or an action, buffer and validate the whole thing.
      How do you move safely to a newer model?
      Offline evals on a set built from real cases, a comparison of cost and latency, then a canary on a few percent of traffic with the same metrics and a rollback ready. The prompt usually needs retuning, because a new model reads the same instructions differently. Test a fallback to a different model the same way, because that is a model change too.

      Sources

      Report an error · Suggest a fix