Home / Chapter 6 · Quality and security
Last edited · 8 min read
Evals
An LLM’s output is random and sensitive to small prompt changes, so “I checked it on three examples” means nothing. An eval is a fixed set of cases with automated grading, run after every change to the prompt, model and tools, plus quality measurement on production traffic.
In plain wordsTests in CI for code that answers slightly differently every time. Instead of “pass or fail” you count how many times out of five it passed, and some assertions can’t be written as string comparisons, so a second model checks them, once you have checked that model.
A revised prompt v2 raises the score from 5 to 7 out of 8. Check in the table what that number hides, then run each case five times and see which change was noise
Illustrative data: a classifier routing support tickets to teams, eight cases. Red rows are cases v2 breaks; greyed-out rows show no meaningful change.
The set and the grading
- Start with error analysis, not metrics: read 50–100 traces, note every problem, group the notes into failure types and count them. The most frequent types show what to fix and what to test. Husain and Shankar put 60–80% of the effort here, not into automated checks.
- Build the set, the golden set, from those failures, not from imagination: turn each type into a few cases with an expected output or a grading criterion. Anthropic advises starting with 20–50 tasks. Add every production failure as a new case, plus edge cases and questions that have no answer.
- With no traffic yet, seed the set with synthetic cases: list the dimensions of a request (user type, intent, difficulty), write a few combinations by hand, have a model expand them and phrase each as a realistic input, then read the traces they produce. Synthetic cases can’t tell you how common a failure is, come out cleaner than real users and miss the quirks of specialised domains and rare languages, so real traces replace them as traffic arrives.
- Not every failure needs an automated evaluator. Fix obvious gaps first, such as an instruction missing from the prompt, and automate only failures that persist: an LLM judge costs time to build and validate, and money on every run.
- Red-teaming is a deliberate adversarial set: people, or an attacker model, try to break the system with prompt injection, jailbreaks, requests outside its policy and attempts to extract data. It is kept apart from quality evals, reported as an attack success rate and run on every release (see “Prompt injection”).
- Grade with the cheapest method that captures the criterion: exact match, a schema or regex, a code check (run it, inspect the state), a model as judge, a human. One criterion per grader and a yes/no verdict instead of a 1–10 scale, because every grader, a model included, applies a scale differently. Skip ready-made metrics such as “helpfulness”, BERTScore or ROUGE: they measure something other than your failures and give false confidence.
- Validate an LLM judge against human labels before trusting its numbers. Raw agreement misleads: if 90% of answers are good, a judge that passes everything agrees 90% of the time and catches no failure. So an expert grades 100–200 examples, you measure separately the share of failures the judge catches (TPR) and of good answers it passes (TNR), and you revise the judge prompt on one part of the labels and check it on another. Known biases: the judge prefers the answer shown first, the longer one and one in its own style (Zheng et al., 2023). Pin the judge’s version and recalibrate whenever it changes.
- Evaluate the system piece by piece and as a whole. In RAG, grade retrieval (recall@k, the share of relevant chunks in the top k results) separately from generation (faithfulness to the chunks, see “RAG”). For an agent, grade the end state rather than a fixed sequence of steps, plus hard constraints on the trace: no forbidden tool calls or policy violations, limits on turns and cost. In a long conversation or agent run, find the first failure: later ones usually follow from it.
Noise and repeated runs
- The same case passes one time and fails the next, even at temperature 0: server-side computation produces slightly different numbers depending on batch size, and batch size changes with load. Run each case several times and report the pass rate.
- A small sample means a lot of noise. With 50 cases and a score around 80%, the 95% confidence interval is about ±11 percentage points, so a 3-point difference between prompts means nothing. Pairwise comparison on the same cases helps, as does looking at the cases whose result changed rather than only at the mean (Miller, “Adding Error Bars to Evals”, 2024).
- Two metrics describe an agent. pass@k: does it complete the task at least once in k attempts, which makes sense when the result can be checked and the best one picked, for example code with tests. pass^k: does it complete it in every one of k attempts, which is the reliability a user expects. At 90% single-attempt success, pass@3 is 99.9% but pass^3 only 73%. In τ-bench (2024), GPT-4o scored below 25% pass^8 on retail customer service.
Offline, CI and production
- Offline you keep two kinds of set. Capability evals start with a low score and show whether a change moves you toward the goal. Regression evals stay near 100% and catch something that used to work and stopped. A set that always scores 100% says nothing about progress, so you add harder cases.
- The regression set is a gate in CI. Run it on every change to the prompt, model, tool definitions and retrieval configuration, because each of them changes behaviour. On a pull request, a fast subset with code-based grading; the full set with an LLM judge nightly or before a release. The gate blocks when a critical case starts failing or the score drops below a threshold with a margin for noise, and it also watches cost and latency.
- In production, a guardrail blocks or fixes an output before the user sees it, so it must be fast and precise; an evaluator measures quality afterwards. There are no expected answers, so you measure differently: a reference-free judge on a sample of traffic (faithfulness to sources, format, safety), behavioural signals (user corrections, retries, escalations to a human, thumbs-up and thumbs-down ratings) and A/B tests or a canary when you change the model. A fixed set run periodically against a pinned version detects changes on the provider’s side (see “Compute and ‘getting dumber’”). Failed production traces flow back into the offline set.
- Public leaderboards tell you which model is worth trying, not whether it will work for you: they measure someone else’s task, have often leaked into training data and saturate quickly.
Check yourself
How do you check that an LLM system works and that a change hasn’t broken it?
Build a fixed set of cases from real failures: read traces, name the error types and turn each into a test. Grade with the cheapest method that suffices: exact match, code checks, an LLM judge validated against expert labels, humans. Run every case several times because outputs are random, and compare versions pairwise. The regression set gates every prompt, model and tool change in CI. In production, grade a sample of traffic with a reference-free judge and watch user signals, and feed failures back into the set.
Po polsku
Zbuduj stały zestaw przypadków z prawdziwych błędów: czytaj trace’y, nazwij typy pomyłek i każdy zamień w test. Oceniaj najtańszą wystarczającą metodą: dopasowanie, test w kodzie, sędzia-LLM sprawdzony na ocenach eksperta, człowiek. Każdy przypadek puszczaj kilka razy, bo wynik jest losowy, a wersje porównuj w parach. Zestaw regresji blokuje w CI każdą zmianę promptu, modelu i narzędzi. Na produkcji oceniaj próbkę ruchu sędzią bez wzorca i patrz na sygnały użytkowników, a nieudane przypadki wracają do zestawu.
Follow-up questions (5)
- How do you start without data?
- With error analysis on 50–100 real examples: read the outputs, name the error types and turn them into test cases. 20–50 cases are enough to start, as long as they come from real failures.
- When can you trust a model as a judge?
- When its grades agree with an expert’s on a held-out sample, measured separately for good and bad answers. A judge that passes everything also scores high accuracy when most answers are good. Give it one criterion and a yes/no verdict, and when comparing two answers, swap their order.
- How do you set up a CI gate that doesn’t block on noise?
- Keep stable regression cases that score close to 100% in the gate, and run each several times. Block when a critical case clearly drops or the mean drops by more than the noise margin. Move cases that flicker without any change into quarantine and investigate them instead of loosening the threshold.
- How do you evaluate an agent?
- With a test of the end state, not the agent’s summary, running each task several times, with pass^k as the reliability metric. Add cost, number of turns and time per task. Traces show where the agent goes astray. You don’t grade the path rigidly, because several routes lead to the goal, but you do check it for forbidden calls and policy violations.
- You’re switching to a cheaper model. How do you decide?
- The same set on both models, pairwise and with repeats, with a separate score for each case type, plus cost and latency. Then a canary or shadow run on part of the traffic and a comparison of production signals. The prompt often needs retuning for the new model, so you compare the best version of each.