# How AI works > An interactive guide for engineers: how language models work under the hood and what matters when building agentic systems. From intuition to nuance, with questions to check yourself. https://howaiworks.dev/ --- ## Tokens *How a model reads and predicts* *Last edited: 28 September 2026* A model never sees letters or words, only a sequence of integers. Each integer is the ID of a piece of text from a fixed vocabulary. Price, context limits and speed are all counted in tokens, and the way text gets split explains several classic model mistakes. **In plain words:** A model reads like someone who knows syllables and whole common words but has never seen individual letters. Familiar words it takes in whole; rare ones it assembles from pieces. *Interactive widget on the page: Pick a tokeniser and an example, or type your own text. See where the token boundaries fall and how many characters one token covers.* ### How the vocabulary is built - The vocabulary is built by BPE (byte-pair encoding). It starts from the 256 possible bytes and, over a large corpus, repeatedly merges the most frequent adjacent pair into a new token until it reaches the target size. Common words end up as a single token; rare ones are assembled from pieces. - Before BPE runs, a regular expression splits the text into words, numbers and punctuation, so merges never cross those boundaries. A space sticks to the word that follows it. In the GPT-4o tokeniser “strawberry” with a leading space is one token, but without the space, e.g. at the start of the text, it is three: st|raw|berry (type the word alone into the widget). - BPE works on UTF-8 bytes, so any text can be encoded and there is no “unknown word” token. Rare combinations fall apart into bytes: in GPT-4o the Polish word “Źdźbło” with a leading space is six tokens, and the space plus “Ź” splits into two pieces, neither of which is a whole character. SentencePiece-based tokenisers get the same guarantee through byte fallback. - Vocabulary size is a design decision: GPT-4 has about 100k tokens, GPT-4o about 200k, Llama 3 128k (100k from the OpenAI tokeniser plus 28k for languages other than English), Gemma 3 262k. A larger vocabulary gives shorter sequences but a bigger embedding table and output layer. - The vocabulary also has special tokens: end of text, role markers, tool calls. They are separate IDs that a well-built API won’t let you produce by typing ordinary text. A conversation is one token sequence with such markers (see “How a model sees a chat”). ### Effects you see in answers - Letters: in “How many r’s are in strawberry?” the model sees one ID instead of ten letters and has to remember the spelling from training. Counting letters, rhyming, reversing words and anagrams are unreliable as a result. Reasoning models do better because they first spell the word out letter by letter, and each letter becomes its own token. - Numbers: GPT-4 and GPT-4o split digits into groups of three from the left (1234567 becomes 123|456|7), Gemma 3 into single digits. With groups, the decimal places of two numbers don’t line up, so “mental” arithmetic is unreliable. For calculations, give the model a tool or a code interpreter. - A prompt that ends with a space breaks the token boundary. In training data a space almost always belongs to the next word, so the model gets a rare pattern and answers worse. Don’t end a raw completion prompt, or a prefilled start of the answer, with a space. This matters mostly when self-hosting: a chat API starts the answer in a new turn, and current Claude models reject prefill altogether (as of September 2026). ### Cost and limits - Price, the context window, `max_tokens`, per-minute rate limits and generation speed are all counted in tokens. Languages other than English use more of them: across the whole text of this guide, the Polish version has about 1.4 times as many tokens as the English one with the GPT-4o tokeniser (3.2 versus 4.5 characters per token) and about 1.6 times as many with Llama 3 (the widget shows the counts). The same content therefore costs more, fills the window faster and takes longer to generate. - The tokeniser belongs to the model. The same text is a different number of tokens across providers, and sometimes across versions of one model: according to Anthropic, the new tokeniser in Claude Opus 4.7 (April 2026) turns the same text into up to about 35% more tokens than Opus 4.6, at an unchanged per-token price. Don’t compare per-million-token prices directly. Count tokens on your own data with the model’s tokeniser or the provider’s endpoint, e.g. `count_tokens` in the Claude API. ### Check yourself **Question:** How does BPE tokenisation work, and how does it affect a model’s cost and behaviour? **Short answer:** A tokeniser splits text into pieces from a fixed vocabulary and maps them to integers. BPE builds the vocabulary from 256 bytes by repeatedly merging the most frequent adjacent pair until it reaches the target size, usually 100–260k. Because it works on bytes, any text can be encoded. Common words become one token; rare and non-English words split into pieces, so Polish takes about 1.3–1.6 times as many tokens as English, depending on the tokeniser. The model never sees individual letters or digits, hence mistakes in letter counting and arithmetic. A larger vocabulary shortens sequences but grows the embedding table and the output layer. ### Follow-up questions - **Why not tokenise by character or by whole word?** Characters make sequences several times longer, and the cost of attention and the KV cache grows with length. Whole words need a gigantic vocabulary, and a typo, a new word or a proper name gets no ID. Byte-level BPE combines short sequences with full coverage. - **What does a larger vocabulary change?** Shorter sequences, so cheaper context and faster generation, especially in languages other than English. The price is a bigger embedding table and output layer (about 1 billion parameters each in Llama 3 70B) and rare tokens that saw few examples in training. - **Can you swap the tokeniser of a trained model?** In practice, no. The embedding table and the output layer are tied to specific IDs, so a new vocabulary needs at least retraining those layers, and usually continued pretraining. Adding a few special tokens during fine-tuning is possible, but their embeddings have to be trained (see “Fine-tuning and LoRA”). - **How would you estimate the cost of a new feature before launch?** On a sample of real data: input and tool definitions with the model’s own tokeniser or the provider’s token-counting endpoint, and output from the usage field of a few real calls. Reasoning models bill thinking tokens as output even when you don’t see them, and no tokeniser can count those in advance (see “Reasoning models”). Price a repeated prefix at the cache rate (see “Prompt caching”), then multiply by the number of calls. Per-million-token prices can’t be compared directly across providers, because the same text is a different number of tokens for each of them. ### Sources - [Sennrich et al.: Neural Machine Translation of Rare Words with Subword Units (BPE, ACL 2016)](https://arxiv.org/abs/1508.07909) - [Anthropic: Introducing Claude Opus 4.7 (new tokeniser, April 2026)](https://www.anthropic.com/news/claude-opus-4-7) - [Andrej Karpathy: Let’s build the GPT Tokenizer](https://www.youtube.com/watch?v=zduSFxRajkE) - [Tiktokenizer: real tokenisers in the browser](https://tiktokenizer.vercel.app/) --- ## Vectors and matrices *How a model reads and predicts* *Last edited: 28 September 2026* Every token becomes a vector of numbers that passes through dozens of layers. A layer is mostly that vector multiplied by large, fixed weight matrices, which is why a model’s cost is counted in parameters, bytes and operations. **In plain words:** Picture a mixing desk with billions of faders. Training set the faders once, and nobody touches them afterwards. Every word passes through the same desk and comes out slightly changed. *Interactive widget on the page: Click a number in the new vector to see which multiplications produced it.* ### A token’s path through the model - The embedding is a plain lookup, not a multiplication: the token ID selects a row of a table. In Llama 3 70B the table is 128,256 × 8192 numbers, over 1 billion parameters. - A layer has two steps. Attention mixes information between tokens (see “Attention”). The MLP processes each token on its own: it expands the vector from 8192 to 28,672 numbers, passes it through a non-linearity (SwiGLU in Llama) and projects it back down. Without the non-linearity, consecutive multiplications could be collapsed into a single matrix. In Llama 3 70B the MLP holds over 80% of a layer’s parameters. - Each step adds its output to the vector instead of replacing it: `x = x + Attn(Norm(x))`, then `x = x + MLP(Norm(x))`. This vector is the residual stream, a shared bus that layers read from and write to. The addition gives the gradient a straight path through dozens of layers, and normalisation (RMSNorm in Llama) keeps the scale of the numbers in check. - At the end, the last token’s vector is multiplied by an 8192 × vocabulary-size matrix. That yields one logit per token, and a softmax turns the logits into probabilities (see “The next token”). *Interactive widget on the page: Click a concept to see its nearest neighbours in vector space.* ### Where the knowledge lives - Inside there is no database of facts and no if-then rules, only numbers in matrices. Research suggests that MLP layers act as key–value memory (Geva et al., 2021): a pattern in the input “lights up” a direction that writes the associated information into the vector. - A fact doesn’t sit in one place. Interpretability research shows that a model packs in more features than it has dimensions by overlapping them, so a single number rarely means one thing. That is why you can’t browse a model’s knowledge or reliably fix a single fact. New or changing knowledge goes into the context (see “RAG”). ### What this means for cost - Memory for the weights is parameters × bytes per parameter: a 70B model takes about 140 GB in BF16 and just under 40 GB at 4 bits (see “Quantisation”). - Compute: about 2 operations (a multiply and an add) per parameter per token at inference, and about 6 in training because of the backward pass. A 70B model needs about 140 GFLOP per token, plus attention, whose cost grows with context and in Llama 3 70B catches up with the weights at about 100k tokens (see “Attention”). In MoE only the active parameters count (see “Mixture of Experts”). - Matrix multiplication is thousands of independent dot products, so it suits GPUs perfectly. Yet when generating a single token, the card mostly waits for the weights to be read from memory (see “Why the GPU is idle”). ### Check yourself **Question:** What happens to a token inside the model, and what does one token cost? **Short answer:** The token ID selects a vector from the embedding table. The vector passes through dozens of blocks: in each, attention mixes in information from earlier tokens and an MLP processes the token on its own. Each step’s output is added to the vector (the residual stream), with normalisation before the step. At the end, a multiplication by the vocabulary matrix yields logits, and a softmax gives the next-token distribution. At short context almost all the cost is multiplication by fixed weight matrices: about 2 operations per parameter per token, so a 70B model needs about 140 GFLOP per token and its BF16 weights take about 140 GB. Attention adds a cost that grows with context length and in a 70B model catches up with the weights at about 100k tokens. ### Follow-up questions - **How much compute and memory does one token cost?** About 2 operations per active parameter at inference and 6 in training, plus attention, which grows with context length. Memory is the weights (parameters × bytes) plus the KV cache, which grows with context length (see “The generation loop and KV cache”). - **Where does a model’s knowledge physically live?** In the weights, largely in the MLP layers, which act as associative memory. Attention moves information between positions. Facts are spread out and overlap one another, which is why it is cheaper to supply new knowledge in context than to write it into the weights (see “RAG” and “Fine-tuning and LoRA”). - **Why residual connections and normalisation?** Adding the output to the input gives the gradient a straight path through dozens of layers, and each layer only has to learn a correction. Normalisation before every step (usually RMSNorm today) keeps the scale of the activations in check. Without them, deep models train unstably. - **How does a token embedding differ from an embedding for search?** A token embedding is a table row, with no context. An embedding model runs a whole passage through a transformer and returns one vector per text, trained so that texts with similar meaning land close together (see “Embeddings and vector search”). ### Sources - [Geva et al.: Transformer Feed-Forward Layers Are Key-Value Memories (2021)](https://arxiv.org/abs/2012.14913) - [Anthropic: Toy Models of Superposition (2022)](https://transformer-circuits.pub/2022/toy_model/index.html) - [Transformer Explainer: a real GPT-2 in the browser](https://poloclub.github.io/transformer-explainer/) - [3Blue1Brown: neural networks and transformers](https://www.3blue1brown.com/topics/neural-networks) - [LLM Visualization in 3D (bbycroft)](https://bbycroft.net/llm) --- ## Attention *How a model reads and predicts* *Last edited: 28 September 2026* Attention is the only place in a transformer where tokens exchange information: each token looks at all the earlier ones and takes a weighted mix of their content. The cost of long context comes down to this mechanism. **In plain words:** Reading “it was hungry”, your eyes jump back to “cat” to know who is meant. Attention is that glance back, done at every word and towards all the previous ones at once. *Interactive widget on the page: Click a token to see which earlier tokens it looks at. The weights in one row always add up to 100%.* ### Mechanism: Q, K, V - Each token projects its vector into three: a query Q (what I’m looking for), a key K (what I can be found by) and a value V (what I pass on). The weights are `softmax(Q·Kᵀ/√d)`: the dot product measures the match, dividing by the square root of the head dimension keeps the softmax from saturating, and the softmax gives weights that sum to 1. The output is a weighted sum of the Vs. In the example, the Q of “it” matches the K of “cat” best, apart from the start token. - Causal mask: a token sees only itself and earlier tokens. This lets training teach prediction at every position of a text at once, and during generation the K and V of old tokens never change, so they can be kept in the KV cache (see “The generation loop and KV cache”). - Multiple heads: in Llama 3 70B the 8192-number vector is projected into 64 query heads of 128 dimensions (and 8 key/value heads). Each head sees the whole vector through its own projection and does its own lookup, and an output matrix combines the results. A few heads can be named, e.g. one that copies a pattern already seen in the text; most can’t. - The attention operation itself ignores order: without the causal mask and without positional information, “2 − 1” and “1 − 2” would look the same. The mask alone lets a model infer positions (a token can count how many predecessors it sees), and Llama 4 has some layers with no positional encoding at all. Most models still add RoPE: it rotates pairs of Q and K coordinates by an angle proportional to position, so the Q·K product depends on the distance between tokens. Context is usually extended after training by rescaling RoPE and a short training run on long texts. *Interactive widget on the page: Lengthen the context and compare three numbers: pairs grow with the square, the KV cache linearly, and attention takes over most of the compute only with a very long prompt.* ### What it costs - Prefill computes n²/2 pairs in every head of every layer. The rest of the layer, the projections and the MLP, grows linearly, though, and dominates at short context: in Llama 3 70B attention only catches up at about 100k tokens. A prompt ten times longer therefore costs 10 to 100 times more compute, depending on its length. - During generation, a new token compares one Q against all the Ks in the cache, so the work per token grows linearly. Memory is what hurts: the KV cache grows linearly with length (43 GB for 128k tokens in Llama 3 70B), and it limits how many conversations fit on a card. - FlashAttention computes exactly the same result, but in tiles in fast on-chip SRAM, without writing the n × n matrix to GPU memory. Memory becomes linear and the computation several times faster, even though the number of operations doesn’t drop (in training it even rises, because the backward pass recomputes attention). - The KV cache is shrunk through architecture: GQA shares one K, V pair across a group of query heads (Llama 3 70B: 8 instead of 64), DeepSeek’s MLA stores one compressed vector, and a sliding window limits some layers to the most recent tokens (Gemma 3: 1,024, gpt-oss: 128). ### Limits of long context - The softmax always hands out 100% of the attention, so a head can’t “look at nothing”. Models learn to dump the excess on the first tokens (attention sinks, Xiao et al., 2023); in the widget above, that is the start token. That is why cutting off the start of the context, e.g. in a naive sliding window, breaks the model, even when the start carried no important content. - Fitting text in the window doesn’t mean the model will use it well. Models use information from the middle of a long context worse than from the beginning and the end (Liu et al., “Lost in the Middle”, 2023). What to put in the context and where is covered in “Context engineering and memory”. ### Check yourself **Question:** How does attention work, and why is long context expensive? **Short answer:** Each token projects its vector into a query, a key and a value. The weights are a softmax over Q·K dot products divided by √d, and the output is a weighted sum of the values. A causal mask hides the future, and many heads do this in parallel. In prefill the number of pairs grows quadratically with length, but in a 70B model attention only dominates compute at around 100k tokens. During generation memory hurts more: the KV cache grows linearly, about 0.33 MB per token in Llama 3 70B (BF16), and limits the batch. FlashAttention, GQA or MLA, and sliding windows help. ### Follow-up questions - **Why divide by √d?** The dot product of vectors with random components has a variance that grows with the dimension. Without scaling, the logits get large, the softmax becomes almost one-hot and the gradients vanish. - **How do MHA, MQA and GQA differ?** In the number of key and value heads. MHA has as many as there are query heads, MQA one shared head, GQA a few groups (Llama 3 70B: 8 for 64 query heads, so an 8 times smaller KV cache). DeepSeek’s MLA stores one compressed vector instead of K and V. Less KV means bigger batches and longer context. MQA and GQA pay for it with a small loss in quality, while DeepSeek reports MLA matching or beating full MHA. - **What does FlashAttention do?** It computes exactly the same attention, but in tiles that fit in the GPU’s fast memory, without writing out the full n × n matrix. Same result, linear instead of quadratic memory and fewer transfers, so it runs faster. It doesn’t reduce the number of operations. - **How do models handle a million tokens?** They combine full attention with cheaper variants, or replace it with them: sliding windows; sparse attention, where each token looks only at the top-k earlier tokens picked by a small, fast scorer (DeepSeek V4 pairs it with a compressed KV cache to serve 1M tokens); linear attention; or Mamba-style layers with a fixed-size state. Add GQA or MLA and rescaled RoPE. Fitting a million tokens is not the same as using them well, so measure long-context quality on your own task. ### Sources - [Vaswani et al.: Attention Is All You Need (2017)](https://arxiv.org/abs/1706.03762) - [Su et al.: RoFormer, the paper that introduced RoPE (2021)](https://arxiv.org/abs/2104.09864) - [Dao et al.: FlashAttention (2022)](https://arxiv.org/abs/2205.14135) - [DeepSeek: V4 release, sparse attention and 1M context (2026)](https://api-docs.deepseek.com/news/news260424/) - [Transformer Explainer: an interactive GPT-2 in your browser (Georgia Tech)](https://poloclub.github.io/transformer-explainer/) --- ## The next token *How a model reads and predicts* *Last edited: 28 September 2026* A model doesn’t return text, only a score (logit) for every token in the vocabulary. Separate code, the sampler, turns the scores into probabilities and draws one token. Every token of an answer is produced this way, one after another. **In plain words:** Your phone keyboard suggests three words. A model does the same for a hundred thousand pieces at once, then rolls a die weighted by those odds. Temperature decides how heavily the die is loaded. *Interactive widget on the page: Change the temperature and top-p, then sample. Compare a question the model knows, one it gets confidently wrong and one where it is guessing.* ### From logits to a token - Softmax with temperature: `p = exp(logit / T) / Σ exp(logit / T)`. T below 1 sharpens the distribution, above 1 flattens it, T → 0 always picks the top token (greedy), and a very high T pushes the distribution towards uniform. Temperature doesn’t change the order of the tokens, only the gaps between them. - Cutting the tail: top-k keeps the k best tokens, top-p (nucleus) the smallest set covering e.g. 90% of the probability, and min-p drops tokens weaker than a fraction of the best one. Top-p and min-p adapt to the model’s confidence: when it is sure, one candidate remains; when it is unsure, dozens. In Hugging Face and vLLM, temperature is applied before truncation, so it affects what survives. - The sampled token is appended to the text and can’t be taken back. Every following token has to fit it, so one unlucky token from the tail can drag a whole made-up sentence behind it. - The sampler can also zero out tokens that don’t fit a JSON schema. That is how structured output works (see “Enforcing output format”). ### Pitfalls - Greedy doesn’t mean best. Always picking the most likely token leads to repetition and loops (Holtzman et al., 2019). DeepSeek recommends a temperature of 0.5–0.7 for R1 precisely to prevent endless repetition. Repetition, frequency and presence penalties lower the logits of tokens already used: they curb loops, but also punish legitimate repeats such as identifiers in code or JSON keys. - Temperature 0 doesn’t give full reproducibility. GPU computations produce slightly different numbers depending on how many requests landed in the batch, and when logits are nearly tied, that changes the chosen token and the entire rest of the answer. The `seed` parameter, where an API has one, doesn’t guarantee determinism. When self-hosting, batch-invariant kernels make the output deterministic at some cost in speed (vLLM: `VLLM_BATCH_INVARIANT=1`). - The distribution doesn’t measure truth. A flat distribution can signal ignorance, but also several equally good phrasings. A model can be confident in an error, and the sampler always picks something, so the sentence sounds just as confident. A lower temperature doesn’t cure fabrication; it only makes it repeatable (see “Hallucinations”). ### What you set in practice - In non-reasoning models: extraction, classification and code at a low temperature, 0–0.3; writing and brainstorming at 0.7–1. Change temperature or top-p, not both at once, and check the effect on an eval set, not on a single sample. - `logprobs` in an API show this distribution, usually before temperature and top-p (the default in vLLM). They are useful for a confidence threshold in classification: you have the model answer with a single label token and read the probability of each label. The Claude API doesn’t return them, and OpenAI does only with reasoning off. - Several samples at T > 0 plus a majority vote (self-consistency) raise accuracy on tasks with one correct answer, and disagreement between the samples signals uncertainty. It costs as many times more tokens as there are samples. - With reasoning models the sampler is often locked (as of September 2026). Claude Opus 4.7 and later, and Claude Sonnet 5, reject non-default `temperature`, `top_p` and `top_k` with a 400 error, and OpenAI with reasoning enabled doesn’t accept `temperature`, `top_p` or `logprobs`. You steer with the reasoning level and the prompt instead. Where a reasoning model does accept them, keep the provider’s recommended values: Google strongly recommends temperature 1.0 for all Gemini 3 models and warns that lower values can cause looping. A low temperature is a tool for non-reasoning models. ### Check yourself **Question:** How does a model choose the next token, and how do you set temperature and top-p? **Short answer:** The model returns logits for the whole vocabulary. A temperature softmax, exp(logit/T), turns them into a distribution: T below 1 sharpens it, above 1 flattens it, and T = 0 always picks the top token. Top-p keeps the smallest set of tokens covering, say, 90% of the probability and cuts the tail where odd tokens come from. On non-reasoning models, use a low temperature for extraction and code and a higher one for creative text, change one parameter at a time and check on evals. T = 0 is not fully reproducible, and reasoning models often lock the sampler or, like Gemini 3, work best at the default 1.0. ### Follow-up questions - **Why doesn’t temperature 0 give full reproducibility?** The result depends on which other requests landed in the same batch: a different batch size means a different order of floating-point operations and slightly different logits. With two nearly tied tokens that is enough to pick the other one, and from there the whole answer diverges. - **Top-k, top-p, min-p: what’s the difference?** Top-k keeps a fixed number of candidates regardless of how confident the model is. Top-p keeps as many as it takes to cover a given share of the probability, so it adapts to the distribution. Min-p drops tokens weaker than a fraction of the best one, e.g. 0.1 × its probability. Its authors report that it copes better with high temperatures, but a 2025 reanalysis found their evidence doesn’t support that. - **Does a lower temperature reduce hallucinations?** Only the ones that come from drawing a token from the tail. If the model “believes” a wrong fact, T = 0 will pick it every time (see “Hallucinations”). - **How do you use logprobs for classification?** Have the model answer with a single label token and read the probability of each label. You get a distribution instead of a single answer, and you set a threshold below which the case goes to a human. Calibrate the threshold on your own data, because after post-training the model’s confidence is often inflated. ### Sources - [Holtzman et al.: The Curious Case of Neural Text Degeneration (top-p, 2019)](https://arxiv.org/abs/1904.09751) - [Hugging Face: How to generate text](https://huggingface.co/blog/how-to-generate) - [Thinking Machines: Defeating Nondeterminism in LLM Inference](https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/) - [Claude docs: migrating to Opus 5.5 (including sampling parameters)](https://platform.claude.com/docs/en/models/opus-5-5/migration-guide) - [Gemini API: Gemini 3 developer guide (temperature)](https://ai.google.dev/gemini-api/docs/gemini-3) --- ## Training *Where a model’s knowledge comes from* *Last edited: 28 September 2026* A model is built in three stages. Pretraining provides language and knowledge through next-token prediction, SFT teaches the assistant role, and RL polishes behaviour and reasoning. After training the weights are frozen, so anything the model doesn’t know has to be given to it in context. **In plain words:** First, years of reading everything in sight. Then a vocational course on example conversations, after which it knows how an assistant behaves. Finally, graded practice: it tries on its own and an examiner rewards good results, so it learns whatever earns points. *Interactive widget on the page: Click a stage and compare what the model does with the same prompt.* ### Pretraining - The task: predict the next token across a huge corpus of text and code (Llama 3: about 15 trillion tokens). The loss is cross-entropy, i.e. −log of the probability the model gave the correct token. Backpropagation computes the gradient, and an optimiser, usually AdamW, nudges every weight slightly towards a smaller error. Every position in the text is a separate example, so one document teaches thousands of predictions at once. - The cost is about 6 × parameters × tokens operations. Chinchilla (2022) found that for a fixed budget the optimum is about 20 tokens per parameter. Today models are trained much longer (Qwen3, 2025: 36 trillion tokens for every size, about 60,000 per parameter for the 0.6B model; Llama 3 8B: almost 2000 per parameter), because a smaller model trained for longer is cheaper at inference over its whole lifetime. - Before training the data is filtered: Llama 3 removed duplicates at the URL, document (MinHash) and line level, dropped junk with heuristics and picked high-quality text with classifiers. Public benchmarks sit on the same web and leak into the corpus. On GSM1k, fresh problems in the style of GSM8K, some models scored up to 8% lower, so a public score can overstate what you will see on your own task (see “Evals”). - Towards the end, the model is trained on progressively longer sequences to extend its context, and on a small portion of the best data, e.g. code and maths, with a decaying learning rate (annealing). This stage is also called mid-training. - The result, the base model, is trained to continue documents, not to answer. Given “How do I bake bread?” it may add more questions from a forum. Pretraining data now includes plenty of question–answer and synthetic instruction text, so modern base models often answer anyway, but unreliably and without the assistant role. ### Post-training: SFT and RL - SFT uses the same loss, but on conversations written in the chat template (see “How a model sees a chat”) and computed only on the assistant’s answer tokens. Hundreds of thousands to millions of examples, written by people and generated by stronger models, teach format, role and tool calling. - RLHF: people compare pairs of answers, a reward model is trained on those comparisons, and the language model is optimised against its score (classically with PPO), with a KL penalty for drifting away from the SFT model. In InstructGPT (2022) a 1.3-billion-parameter model trained this way beat the 175-billion-parameter GPT-3 in human ratings. DPO learns directly from better/worse pairs, without a separate reward model. Today the comparisons are often made by a model judge following written principles (RLAIF, Constitutional AI); in Tülu 3, GPT-4o rated the answers. - RLVR: the reward is checked by a program, such as tests or a comparison with the correct result. The model generates many attempts, and the better ones are reinforced. GRPO compares attempts on the same prompt with each other, without a separate critic. This is how reasoning models came about: a long chain of thought pays off because it raises the chance of a reward (see “Reasoning models”). - RL in environments: the same idea for agents. Multi-turn tasks with tools (a terminal, a repository with tests, simulated APIs) are rewarded for the end state; Kimi K2 (2025) had a joint RL stage in real and synthetic environments. SFT teaches the tool-call format, and this is where the model practises long, multi-step runs. - Distillation: a smaller model learns from the answers or probability distributions of a larger one, which is why small models beat what their size suggests (see “Distillation”). ### What this means in practice - After training the weights are frozen. The model doesn’t learn from the conversation, and “memory” in apps is notes appended to the context (see “Context engineering and memory”). About the world after its cutoff date it knows only what it gets in context. - RL optimises the reward, not the intent (reward hacking). A model rewarded for passing tests can learn to game them instead of fixing the code, and one rewarded by human ratings learns to agree with people (sycophancy). In RLHF a KL penalty and diverse rewards limit this but don’t eliminate it. RLVR for reasoning often drops the KL term (DAPO, 2025), because a long chain of thought has to move far from the starting model, so there the quality of the verifier is what holds reward hacking back. - Your own training starts from an existing model. Fine-tuning is good at teaching style, format and narrow tasks, and poor at teaching new knowledge. Knowledge that changes or needs a source goes into the context (see “Fine-tuning and LoRA” and “RAG”). ### Check yourself **Question:** How is an LLM made: what is the difference between pretraining, SFT, RLHF and RLVR? **Short answer:** Pretraining teaches next-token prediction on trillions of tokens of text with a cross-entropy loss. That gives language and knowledge, but the base model only continues documents. SFT on conversations in the chat template teaches the assistant role and the tool format. RLHF optimises against a reward model learned from comparisons made by people or a model judge, with a KL penalty for drifting away from the SFT model. RLVR optimises against a reward checked by a program, such as tests: that is how reasoning models are made, and, in multi-turn tool environments, models for agents. The risk of RL is reward hacking. After training the weights are frozen, so current knowledge has to be supplied in context. ### Follow-up questions - **RLHF versus RLVR?** RLHF takes its reward from a model trained on the preferences of people or a model judge: good for style and helpfulness, but the reward model can be fooled. RLVR takes its reward from an automatic check, such as tests or the task’s result: harder to fool and easy to scale, but only where a program can verify the result. - **PPO, DPO, GRPO?** PPO is classic RL with a reward model and a separate critic that estimates the value of a state. DPO learns directly from “better/worse answer” pairs, with no reward model and no generation in the loop. GRPO (DeepSeek) generates several answers to the same prompt and compares them with each other, so it needs no critic. - **Why a KL penalty in RL?** It keeps the model close to the SFT model. Without it, optimisation quickly finds answers that the reward model rates highly and people rate poorly, and the model loses general skills. RLVR for reasoning often drops it (DAPO), because the model has to move far from its starting point; there a good verifier is what stops the model gaming the reward. - **When fine-tuning and when RAG?** Fine-tuning for style, format and narrow skills. RAG for knowledge that changes and needs a source (see “Fine-tuning and LoRA” and “RAG”). ### Sources - [Ouyang et al.: Training language models to follow instructions with human feedback (InstructGPT, 2022)](https://arxiv.org/abs/2203.02155) - [Lambert et al.: Tülu 3, an open post-training recipe with RLVR (2024)](https://arxiv.org/abs/2411.15124) - [Llama Team: The Llama 3 Herd of Models (2024), data and decontamination](https://arxiv.org/abs/2407.21783) - [Kimi Team: Kimi K2: Open Agentic Intelligence (2025)](https://arxiv.org/abs/2507.20534) - [Andrej Karpathy: Deep Dive into LLMs like ChatGPT](https://www.youtube.com/watch?v=7xTGNNLPyMI) --- ## Distillation *Where a model’s knowledge comes from* *Last edited: 28 September 2026* Distillation trains a small model, the student, on the outputs of a large one, the teacher. It is why small models beat what their size suggests and how reasoning reached 7B models, and it lets you replace an expensive model on one task with a cheaper one. **In plain words:** A chess student who sees only a grandmaster’s moves learns more slowly than one who is also told which other moves were nearly as good and which lose at once. On-policy distillation is the grandmaster watching the student’s own games and marking every move, instead of showing their own. The student rarely surpasses the master, but does pick up the master’s mistakes. *Interactive widget on the page: Switch the target the student learns from and change the temperature. Compare what one label says with what the teacher’s distribution says.* ### Hard labels and distributions - Hard labels (sequence-level distillation): the teacher generates answers to your prompts, and the student is trained on them with ordinary SFT (see “Training”). Text is all you need, so it works through an API and across different tokenisers. This is the most common form of distillation. - Distributions (logit distillation): at every position the student minimises the KL divergence from the teacher’s full next-token distribution to its own. Hinton et al. (2015) raise the softmax temperature in both models (see “The next token”) to bring out small probabilities, and multiply this part of the loss by T², because its gradients shrink as 1/T². Which wrong answer is almost right lives in the ratios of very small probabilities; Hinton called this “dark knowledge”. - A distribution carries more information per example: a ranking of alternatives instead of a single index, and a gradient that varies less between examples. In the paper’s speech experiment, on 3% of the data hard labels overfitted at 44.5% accuracy, while the teacher’s distributions reached 57.0%, against 58.9% for training on all the data. - The price: you need the teacher’s logits, so open weights or your own model. An API returns at most a short list of logprobs, and Claude, or OpenAI with reasoning on, none at all. Both models need the same tokeniser, because the distribution is over token IDs. A full distribution is 100–260k numbers per position, so Gemma 3 stores 256 tokens sampled by teacher probability, and zeroes and renormalises the rest. ### On-policy distillation - Both methods above train on the teacher’s text. When generating, the student makes mistakes the teacher never makes, lands in situations it never saw in training, and the errors compound. - In on-policy distillation the student generates the answer and the teacher, in one forward pass, computes the probability of each of its tokens. The loss is a per-token reverse KL: the less likely a token is for the teacher, the bigger the penalty. It works like RL, but with every token graded instead of one reward for the whole answer. Reverse KL pushes the student towards one of the teacher’s ways of answering instead of spreading across several. - Qwen3 (2025) trained its 0.6B to 14B and 30B-A3B models this way, with Qwen3-32B or 235B-A22B as the teacher: first SFT on the teacher’s answers, then an on-policy stage. On Qwen3-8B that stage reached 74.4% on AIME’24 in 1,800 GPU hours, while RL from the same starting point reached 67.6% in 17,920 hours. Thinking Machines spelled out the method in “On-Policy Distillation” (October 2025). ### In the pretraining of small models - Gemma 2 (2024) trained its 2B and 9B models on a larger teacher’s distributions, on more than 50 times the compute-optimal number of tokens. In an ablation, a 2B model after 500B tokens averaged 67.7 points with distillation from a 7B model and 60.3 without it. In Gemma 3 (2025) every size is distilled. - Llama 3.2 1B and 3B (September 2024) were pruned from Llama 3.1 8B, and during pretraining the logits of the 8B and 70B models served as targets for every token. - The usual pretraining target is the one token a document happened to contain; a teacher’s distribution says what could have been there, so every token teaches more. That is one reason 1–4B models beat their size. ### Reasoning distillation - DeepSeek-R1 (January 2025): about 800k examples, including about 600k reasoning traces kept only when their result was correct. SFT alone on them, with no RL, produced models from Qwen2.5 1.5B to Llama 3.3 70B. Qwen2.5-32B scored 72.6% on AIME 2024 after distillation, while the same base after more than 10k steps of large-scale RL scored 47.0%. - A small, carefully chosen set also works. s1 (2025): 1,000 questions with reasoning traces from Gemini Flash Thinking, and 26 minutes of training Qwen2.5-32B on 16 H100s. LIMO (2025): 800 selected solutions and 63.3% on AIME24. The knowledge is already in the base; the traces teach the form of long reasoning (see “Reasoning models”). ### Limits - The student is capped by the teacher and copies its errors and biases; the R1 authors note that going further needs a stronger base and more RL. Hard labels even lock in guesses: if the teacher sampled a name it didn’t know, the student learns to state it with confidence. - Too big a gap between teacher and student also hurts. In Gemma 3’s ablation the smaller teacher won with short training and the larger one only with long training. Li et al. (2025): models up to 3B don’t consistently gain from strong teachers’ long reasoning traces and learn better from a mix of short and long ones. - Style transfers more easily than skill. Gudibande et al. (2023): raters judged models trained on ChatGPT’s answers competitive, but targeted tests showed they closed almost none of the gap outside tasks well covered in the data. Scores on benchmarks close to the distillation data transfer; robustness to unusual inputs, much less. ### The labs’ side - Terms of service forbid using outputs to train competing models (as of September 2026). OpenAI Terms of Use (effective 1 January 2026): “Use Output to develop models that compete with OpenAI.” The API’s Services Agreement (same date) makes two exceptions: undistributed models that categorise, classify or organise data, such as embeddings and classifiers, and fine-tuning OpenAI’s own models. Anthropic Commercial Terms (effective 17 June 2025): customers may not “access the Services to build a competing product or service, including to train competing AI models” unless Anthropic expressly approves. Google Gemini API terms (effective 23 March 2026): “You may not use the Services to develop models that compete with the Services.” - The raw chain of thought is hidden. For o1 (September 2024) OpenAI gave user experience, competitive advantage and the option to monitor the chain of thought as its reasons. Claude’s documentation says summarising “prevents misuse” and that no setting returns the raw chain of thought. Gemini returns summaries without a stated reason. - Anthropic (23 February 2026) says it detected campaigns by DeepSeek, Moonshot AI and MiniMax: over 16 million exchanges through about 24,000 fraudulent accounts, aimed at reasoning, tool use and coding, including requests to write out, step by step, the reasoning behind a finished answer. These are one party’s claims. ### Distillation in your system - The typical case: a large model with a long prompt does one task well but is too expensive or too slow at your volume. Collect a few thousand real inputs, generate answers with the teacher, discard bad ones with a validator or a judge, and fine-tune a small model, for example with LoRA (see “Fine-tuning and LoRA”). The student no longer needs the long prompt, so you pay for fewer input tokens and get answers faster. Compare student candidates as in “Choosing a model”. - Distributions and on-policy distillation become options when the teacher has open weights and the same tokeniser, such as a larger model from the same family. OpenAI explicitly allows fine-tuning its own models and building internal classifiers, but none of these terms says whether a narrow generative model for your own use “competes”, so that is a question for a lawyer. An open teacher with a suitable licence avoids it: DeepSeek-R1’s MIT licence names distillation explicitly. - Test the student on an eval set of real cases, set aside before you generate the data (see “Evals”): side by side with the teacher, with unusual cases and refusals included. Cases where the student loses can be routed to the teacher. ### Check yourself **Question:** How does distillation move a large model’s ability into a small one, and how do hard labels, teacher distributions and on-policy distillation differ? **Short answer:** The student, a small model, learns to imitate the teacher, a large one. Most often with hard labels: the teacher generates answers and the student is trained on them with ordinary SFT, so text from an API is enough. Logit distillation minimises the KL divergence between the teacher’s and the student’s full next-token distributions, usually with a temperature. Each example then carries more information, because it also shows which alternatives are close, but it needs the teacher’s logits and a shared tokeniser. In on-policy distillation the student generates and the teacher grades each of its tokens with reverse KL, so the student also learns from its own mistakes; in Qwen3 this beat RL at about a tenth of the GPU hours. The student won’t surpass the teacher, copies its errors and picks up style more easily than skill, so test it on your own eval set. The terms of OpenAI, Anthropic and Google forbid using outputs to train competing models. ### Follow-up questions - **If distributions carry more information, why is most distillation done with hard labels?** Because hard labels need only text. A distribution needs the teacher’s logits, so open weights, and an API returns at most a short list of logprobs, or none. Both models need the same tokeniser, and a full distribution is 100–260k numbers per position, so a sample of tokens is stored instead, such as Gemma 3’s 256. With a teacher behind an API, SFT on its answers is what remains. - **Why does on-policy distillation beat SFT on the same teacher’s answers?** SFT trains in situations the teacher reaches. When generating, the student makes its own mistakes, lands in situations it never saw, and the errors compound. On-policy distillation trains on the student’s own text while the teacher grades every token, so the signal is dense like SFT and comes from the student’s distribution like RL. The price is a teacher forward pass on every student sample and access to its logprobs. - **The student matches the teacher on the eval set, but in production it fails on slightly different inputs. What happened?** The distillation data covered the eval set’s distribution rather than production traffic, and the student picked up the teacher’s style more easily than its skill. What helps: broadening the inputs with real production cases, including hard and unusual ones, a fresh eval set drawn from production, and routing to the teacher the cases where a validator rejects the student’s result. - **May you distil a model available only through an API into your own model?** Technically, hard labels are enough. The terms (as of September 2026) forbid using outputs to train competing models. OpenAI makes exceptions for undistributed classifiers and embeddings and for fine-tuning its own models, and Anthropic allows it only with express approval. The terms don’t say whether a narrow model for your own use “competes”, so that is a question for a lawyer. An open teacher whose licence allows distillation, such as DeepSeek-R1 under MIT, avoids the question. - **Does a bigger teacher always make a better student?** No. In Gemma 3’s ablation the smaller teacher won with short training and the larger one only with long training. Models up to 3B don’t consistently gain from strong teachers’ long reasoning traces, and a mix of short and long traces helps them (Li et al., 2025). Pick the teacher by the student’s score on the eval, not by the teacher’s own score. ### Sources - [Hinton, Vinyals, Dean: Distilling the Knowledge in a Neural Network (2015)](https://arxiv.org/abs/1503.02531) - [DeepSeek-AI: DeepSeek-R1 (2025), section on distillation into Qwen and Llama](https://arxiv.org/abs/2501.12948) - [Qwen Team: Qwen3 Technical Report (2025), strong-to-weak distillation](https://arxiv.org/abs/2505.09388) - [Thinking Machines: On-Policy Distillation (October 2025)](https://thinkingmachines.ai/blog/on-policy-distillation/) - [Anthropic: Detecting and preventing distillation attacks (February 2026)](https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks) --- ## Fine-tuning and LoRA *Where a model’s knowledge comes from* *Last edited: 28 September 2026* Fine-tuning is further training of an existing model on your examples: it is good at changing how the model answers and poor at changing what it knows. LoRA does it cheaply by freezing the model and training a small correction, often under 1% of the weights. **In plain words:** Fine-tuning is on-the-job training for an experienced employee. Afterwards they write in the company’s format and tone, but they won’t memorise the price list; that’s what the binder on the desk is for, i.e. RAG. LoRA is corrections on a transparent sheet laid over the textbook: the original stays untouched, and the sheet can be lifted off or swapped for another. *Interactive widget on the page: Pick a problem and see which rung to start from.* ### What it changes and what it doesn’t - Fine-tuning shifts the weights so that answers resemble your examples. It is best at teaching what shows up in every answer: format, style and tone, consistent behaviour (when to ask a clarifying question, when to refuse) and a narrow skill such as classification or extracting fields from a document. It also transfers a large model’s behaviour on one task to a smaller, cheaper and faster one, trained on the large model’s answers (see “Distillation”). - Facts go into the weights poorly and with risk. Gekhman et al. (2024): examples with knowledge the model didn’t have are learned more slowly than the rest, and once learned, they linearly increase the tendency to hallucinate. Ovadia et al. (2023): RAG beat continued training on raw text, for both known and new knowledge. A fact in the weights also has no source and can’t be corrected without another training run. - The order: prompt and context (instructions, examples, structured output), then RAG for knowledge (see “RAG”), fine-tuning last. You change a prompt in a minute and update RAG by adding a document, while fine-tuning means data, training, evals and a new model to maintain. Reach for it when a polished prompt still loses on evals, or when a long prompt full of examples is too expensive at your volume. ### Types and where to train - SFT (supervised fine-tuning): input → reference answer pairs. The loss is computed only on the answer tokens, as in the SFT stage described in “Training”, but on your examples. The model imitates the references, so their quality sets the ceiling. - Preference tuning, e.g. DPO (Rafailov et al., 2023): pairs of a better and a worse answer to the same prompt. The model raises the probability of the better one relative to the worse one, with no separate reward model and no RL loop. It suits tone and things to avoid, because a difference is easier to show than a perfect reference is to write. - Reinforcement fine-tuning: the model generates its own answers and a grader scores them. The grader is a function in code, a comparison with the expected result, or a judge model. No reference answers are needed, only a way to score, so it suits tasks with a verifiable result. The model will exploit a leaky grader (reward hacking, see “Training”). - Hosted fine-tuning of closed models: you upload a JSONL file of examples, the provider trains and serves the model, and you don’t get the weights. Google Vertex AI offers SFT, preference tuning and RL for Gemini models (September 2026). OpenAI is winding down its platform: since 7 May 2026 it hasn’t accepted organisations that never fine-tuned, since 2 July 2026 it also blocks those with no inference on a fine-tuned model in the last 60 days, and from 6 January 2027 no one will be able to create a new training job. - An open-weight model (see “Open-weight models”) can be trained through a managed service: Together AI takes a file of examples, and Thinking Machines’ Tinker gives you an API for your own SFT, DPO or RL loop on LoRA. They run the GPUs, and you download the adapter and serve it anywhere. Training on your own hardware gives full control of the pipeline, but the GPUs and serving are on you. ### How LoRA works - Full fine-tuning updates every weight. LoRA (Hu et al., 2021) freezes a d × k matrix W and learns only a correction ΔW = B·A, where B is d × r and A is r × k. The rank r (e.g. 8, 16, 64) is much smaller than d and k, so instead of d·k parameters you train r·(d + k): for a 4096 × 4096 matrix and r = 16 that is 131k instead of 16.8 million. The authors assume that the weight change needed to adapt a model has low rank. B starts at zero, so at first ΔW = 0 and the model behaves like the base, and the output is scaled by α/r. - The original paper added LoRA mainly to attention. Schulman (Thinking Machines, 2025) shows that LoRA on all layers, especially the MLP, matches full fine-tuning on small and medium datasets, and in RL even at rank 1. It loses when there is more data than the adapter can hold, and it tolerates large batches worse than full fine-tuning. The optimal learning rate is about 10 times higher than for full fine-tuning. - QLoRA (Dettmers et al., 2023) is LoRA on a base quantised to 4 bits in the NF4 format (see “Quantisation”). Gradients flow through the frozen base into a 16-bit adapter. The base takes about 4 times less memory, and in the paper fine-tuning a 65B model fitted on a single 48 GB card with no loss of quality compared with 16-bit training. *Interactive widget on the page: Change the rank and the model class. See how many weights LoRA trains and how much GPU memory it needs.* - An adapter is a small file: for an 8B-class model with r = 16, about 42 million parameters, i.e. about 84 MB in BF16 against a 16 GB base. You keep one base and many adapters, e.g. per customer, language or task. vLLM keeps one copy of the base and mixes requests for different adapters in one batch; S-LoRA (Sheng et al., 2023) serves thousands of adapters on a single GPU this way. The price is an extra B·(A·x) multiplication at every step. - Merging: after training, B·A can be added to W. The model is then exactly as fast as the base, but every variant is a separate full copy. Several adapters on the same base can also be combined into one without training, by a weighted sum or with TIES and DARE (`add_weighted_adapter` in Hugging Face PEFT). Skills can interfere with each other in the process, so evaluate a merged adapter like a new model. ### Data, evaluation, maintenance - Data quality beats quantity. In LIMA (Zhou et al., 2023), 1000 carefully chosen SFT examples were enough for a 65B model to give high-quality answers. The authors conclude that knowledge comes from pretraining and tuning mostly teaches form. The model learns everything that repeats in the data, including typos, inconsistent formatting and bad answers. - Set the eval set aside before training and never train on it (see “Evals”). First measure the base with your best prompt on it: that is the bar fine-tuning has to clear. Add out-of-task cases and refusal tests, because fine-tuning weakens safeguards: Qi et al. (2023) stripped them from GPT-3.5 Turbo with 10 examples for less than $0.20, and ordinary, benign datasets weakened them too, just less. - Overfitting: training loss falls, validation loss rises, and the model parrots phrases from the examples. Fewer epochs, a lower learning rate and picking the checkpoint by validation loss help. Catastrophic forgetting: the model loses skills outside your data. LoRA forgets less than full fine-tuning, but on large datasets it also learns less (Biderman et al., 2024, on code and maths: the capacity-limited case above). - A fine-tuned model is tied to its base. An adapter fits only that exact version of the weights, and with a hosted provider the model disappears along with its base: on 23 October 2026 OpenAI shuts down ft-gpt-4.1-nano, among others. A new base means a new training run, so you version the data, configuration, base version, adapter and eval results together, and run the whole pipeline with one command. With every new base, first check whether a prompt alone is now enough. ### Check yourself **Question:** When would you use fine-tuning instead of a prompt or RAG, and how does LoRA work? **Short answer:** Order: prompt and context first, RAG for knowledge, fine-tuning last. Fine-tuning is good at behaviour: format, tone, a narrow task, or distilling a large model’s behaviour into a small one. It is poor at facts: they go in unreliably, without a source, every change means retraining, and they can increase confident errors. LoRA freezes the weights and learns a low-rank update ΔW = B·A of rank r, so r·(d + k) parameters instead of d·k. An adapter is tens of MB, so many adapters can be served on one base. QLoRA does the same on a 4-bit base. ### Follow-up questions - **You have 50k company documents that change every week. Fine-tuning or RAG?** RAG. Facts from fine-tuning go in unreliably, have no source, and every change needs a new training run. Fine-tuning can come later to teach the model the answer format and how to cite passages, but the knowledge stays in the index. - **How do you choose the LoRA rank and the layers to apply it to?** All linear matrices, including the MLP, because attention-only LoRA clearly lags behind. Rank is the adapter’s capacity: on small and medium datasets a low rank is enough, in RL even r = 1, while on large datasets LoRA starts losing to full fine-tuning. The learning rate is about 10 times higher than for full fine-tuning; compare a few ranks on the eval. - **You have 200 customers and each wants a model in their own style. How do you serve that?** One base and one LoRA adapter per customer, without merging. vLLM picks the adapter for each request, keeps several active adapters in one batch next to a single copy of the base and loads more on the fly. The price is a little extra compute at every step instead of 200 full copies of the model. - **After fine-tuning the format is perfect, but the model does worse outside the task and is easier to talk into things it used to refuse. What happened?** Catastrophic forgetting and weakened safeguards: training pushed the weights only towards your examples. Fewer epochs, a lower learning rate or LoRA instead of full fine-tuning help, as do general examples and refusals in the data, with an out-of-task set and safety tests in the eval. - **The provider retires the base model your fine-tune is built on. What do you do?** First check whether the new base with just a prompt already passes the eval, because then fine-tuning is no longer needed. If not, retrain on the new base with the same pipeline. Data, configuration and the eval set are versioned, so it is a single run. ### Sources - [Hu et al.: LoRA: Low-Rank Adaptation of Large Language Models (2021)](https://arxiv.org/abs/2106.09685) - [Dettmers et al.: QLoRA: Efficient Finetuning of Quantized LLMs (2023)](https://arxiv.org/abs/2305.14314) - [John Schulman, Thinking Machines: LoRA Without Regret (2025)](https://thinkingmachines.ai/blog/lora/) - [Gekhman et al.: Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations? (2024)](https://arxiv.org/abs/2405.05904) - [Thinking Machines: Tinker, an API for training open-weight models with LoRA](https://thinkingmachines.ai/tinker/) --- ## How a model sees a chat *Where a model’s knowledge comes from* *Last edited: 28 September 2026* The model never sees a chat window. The system prompt, messages, tool calls and tool results are concatenated into one document with role markers, and the model continues the text after the assistant marker. This explains the cost of long conversations, the limits of the system prompt and most bugs in self-hosted deployments. **In plain words:** A script where every line has a speaker label: DIRECTOR, CUSTOMER, ASSISTANT. An actor with no memory gets the whole script from page one every time, sees an empty ASSISTANT label at the end and writes the next line. If someone slips a stranger’s page into the script, the actor reads it just like the rest. *Interactive widget on the page: Add messages and tool calls, and switch views. Notice that every request resends the whole document from the start.* ### The chat template - Role markers are special tokens: separate IDs in the vocabulary, not ordinary characters (see “Tokens”). The model learned the format during SFT (see “Training”): the assistant marker is followed by the answer, and the answer by an end-of-turn token. - Every model family has its own template: ChatML in Qwen (`<|im_start|>user … <|im_end|>`), Llama 3 (`<|start_header_id|>user<|end_header_id|> … <|eot_id|>`), harmony in gpt-oss (`<|start|>user<|message|> … <|end|>`), Mistral (`[INST] … [/INST]`). You send a list of messages with roles, and the server renders it with the template the model was trained on. - In Hugging Face the template is a Jinja template shipped with the tokeniser. `apply_chat_template(messages, add_generation_prompt=True)` builds the document and appends the assistant marker at the end. Without that marker the model may carry on writing the user’s message instead of answering it. - Generation stops at the end-of-turn token, provided the server has it in its list of stop tokens. In Llama 3 a turn ends with `<|eot_id|>`, not `<|end_of_text|>`. A server that only waits for the latter lets the model write a user marker and invent the user’s next message. ### The whole document every turn - The model doesn’t remember the conversation. Every request carries the whole document from the start, so the total number of tokens sent grows with the square of the number of turns (see “The context window and agents”). Server-side state, such as OpenAI’s `previous_response_id`, only saves the upload: the server rebuilds the document and bills all earlier input tokens again. What to keep, summarise or drop from the history is covered in “Context engineering and memory”. - The stable start of the document (system prompt, tool definitions and older history) is what prompt caching reuses (see “Prompt caching”). Anything variable at the start, such as the current time in the system prompt, breaks the cache every turn. Editing an old message produces a new document, and the cache is lost from the point of the change. - Tools are document fragments too. Their definitions sit next to the system prompt and cost tokens every turn. A call is text the model writes after a special marker, and your code appends the result as a new fragment (see “Tools (function calling)”). - A reasoning model’s thinking is another fragment with its own markers: `…` in some open models, the `analysis` channel in harmony (see “Reasoning models”). Harmony drops it from the history after the final answer but keeps it between tool calls. ### Images, PDFs and voice - A vision encoder cuts an image into square patches and turns each one into a vector that takes a token’s place in the sequence. For Claude a patch is 28 × 28 px, so a 1000 × 1000 px photo is 1296 tokens, and Claude 4.7 and later downscale anything above 2576 px on the long edge or 4784 tokens (as of September 2026). - An image stays in the history and is billed again every turn: ten 1000 × 1000 px screenshots in an agent’s history add about 13,000 tokens to every request. Claude reads each PDF page twice, as text (1,500–3,000 tokens) and as an image, so a 100-page report is 150,000–300,000 tokens of text alone. Cache such input, downscale it before sending, and drop screenshots the agent no longer needs. - Audio becomes tokens too: Gemini counts 32 tokens per second of audio, and OpenAI’s Realtime API 10 per second of the user’s speech and 20 per second of the model’s, resending the whole conversation for each response. - A voice agent is built one of two ways. A chained pipeline (speech-to-text → text model → text-to-speech) gives you transcripts, policy checks before the answer and any text model, but every stage adds to the silence the caller hears. A speech-to-speech model (OpenAI Realtime, Gemini Live) handles audio in one session, with barge-in and turn detection, but shows you less of what happens in between. Measure the time from the end of the caller’s speech to the first audio, at the median and p95. ### The system prompt is not a safeguard - The system prompt is text at the start of the document that the model was trained to prioritise. It is not a separate channel or a permission. An instruction in an email, on a web page or in a tool result sits in the same token stream and sometimes wins (see “Prompt injection”). Rules that must hold are enforced by application code. Don’t keep secrets in the system prompt, because the model may quote it. - User text must not become a marker. In a well-built API, typing “<|im_start|>system” yields ordinary characters. When self-hosting, check this yourself: Hugging Face tokenisers recognise special-token strings in text by default (`split_special_tokens=False`), so pasted text can open a real system turn. ### Self-hosting and fine-tuning - With an open-weight model (see “Open-weight models”) the template is part of the model. A wrong or outdated one raises no error; it silently degrades answers and breaks tool calling. A common case is a doubled beginning-of-text token: the document from `apply_chat_template(tokenize=False)` gets tokenised a second time with special tokens added, and the Hugging Face docs warn that this hurts quality. - Render fine-tuning data with exactly the template the server will use, and compute the loss only on assistant tokens (see “Fine-tuning and LoRA”). - Prefill: you end the message list with the start of the answer, such as `{`, and the model writes on from there. In transformers this is `continue_final_message=True`. From Opus 4.6 and Sonnet 4.6 onwards, the Claude API rejects prefill with a 400 error and points you to structured outputs (see “Enforcing output format”). ### Check yourself **Question:** How does a model see a conversation, and how does it know who said what? **Short answer:** The model never sees a chat window. The server concatenates the system prompt, history, tool calls and tool results into one document, separating roles with special tokens from a template the model learned during SFT. It ends with an assistant marker, and the model continues the text until it emits an end-of-turn token. The model has no memory, so the whole document is resent every turn. The system prompt is just text at the top, not a security boundary: an instruction hidden in an email sits in the same stream. ### Follow-up questions - **Why does the model sometimes carry on writing as the user?** Because it didn’t generate the end-of-turn token, or the server doesn’t have that token in its stop list. To the model it is still one document, so the most likely continuation is a user marker followed by the user’s next message. A wrong template or stop configuration is usually to blame. - **Can the model tell an instruction in the system prompt from one in an email?** Only as far as it learned to in training. Models are trained on a hierarchy: system above user, user above tool content (OpenAI, “The Instruction Hierarchy”, 2024). This improves robustness but is not a hard boundary, because everything sits in one token stream. - **Why does the API take a list of messages instead of ready-made text?** So that the template, markers and tool format always match what the model was trained on, and so that user content can’t forge a role marker. - **You are deploying an open-weight model on your own server. What do you check in the template?** That the tokeniser’s template matches the model card, that the end-of-turn token is in the stop list, that there is no doubled beginning-of-text token, and that user text can’t turn into special tokens. The simplest check is to render a sample conversation with tools and compare it character by character with the example in the model’s documentation. - **What happens when you edit an old message?** The model remembered nothing, so it simply gets a new document. You lose the prompt cache from the point of the change to the end. On Claude Opus 5.5 and Fable 5.1, editing history before a thinking block ends in a 400 error (see “Reasoning models”). ### Sources - [Hugging Face: Chat templates](https://huggingface.co/docs/transformers/main/en/chat_templating) - [OpenAI: the harmony format (gpt-oss)](https://developers.openai.com/cookbook/articles/openai-harmony) - [Wallace et al.: The Instruction Hierarchy (2024)](https://arxiv.org/abs/2404.13208) - [Claude docs: Vision (image tokens)](https://platform.claude.com/docs/en/build-with-claude/vision) - [OpenAI: Voice agents](https://developers.openai.com/api/docs/guides/voice-agents) --- ## Reasoning models *Where a model’s knowledge comes from* *Last edited: 28 September 2026* A reasoning model writes a chain of thought before it answers: it tries, checks and corrects itself. It is the same token-by-token machinery, trained with RL on tasks with verifiable outcomes, and you pay for thinking tokens as output. **In plain words:** A student allowed to work on scratch paper during a test makes fewer mistakes than one who writes the answer straight down. But scratch paper won’t help them recall a date they never learned. Working it out takes time, and for a model every line of it goes on the bill. *Interactive widget on the page: Pick a task and a thinking level. See where thinking improves the result and where it only adds cost.* ### Where thinking comes from - There is no separate thinking module. The chain of thought is ordinary tokens from the same loop as the answer (see “The next token”), just marked as thinking. Every step written down stays in the context, so later tokens build on it like on scratch paper. - The model learns to write this scratch paper through RL (see “Training”) on automatically checked tasks: maths with a known answer, code with tests. The reward goes mainly to the result, so whatever leads to it gets reinforced, including self-checking and backtracking. In DeepSeek-R1-Zero these behaviours became frequent through RL alone, without human-written examples, though base models already show them occasionally (Liu et al., 2025). RLVR mostly makes the model reach, more reliably, answers the base model could already find: given enough samples (large pass@k), the base model catches up (Yue et al., 2025). - It is one form of test-time compute: spending more computation at answer time instead of building a bigger model. Others are sampling several attempts and voting, or picking the best one with a verifier. Gains shrink as thinking gets longer, and on some tasks longer thinking makes results worse (inverse scaling in test-time compute, 2025). ### When it pays off - It helps on multi-step tasks where mistakes come easily: calculations, code, planning, analysis with many conditions. On simple facts, lookups and short classification it only adds cost and waiting. - It adds no knowledge. What the model doesn’t know it won’t work out: you get a longer, more confident-sounding guess (see “Hallucinations”). - Every thinking token is a decode step that waits on memory (see “Why the GPU is idle”). A few thousand thinking tokens means tens of seconds before the first word of the answer. In a UI, show progress or a summary of the thinking. Run tasks with no human in the loop in the background or as batch jobs. ### Controls and the bill - You control the effort level (`reasoning.effort` in OpenAI, `effort` in Anthropic, `thinking_level` in Gemini, from `none` to `max` depending on the model). A fixed thinking-token budget (`budget_tokens`, `thinking_budget`) survives only on older models: Claude 4.7 and later reject it with a 400, and Gemini 3 accepts it only for backward compatibility. Even there it is a target, not an allocation. You pick the level with evals on your own tasks (see “Evals”). - On the newest flagship models thinking can’t be switched off (as of September 2026): Claude Opus 5.5 and Fable 5.1 reject `thinking: {type: "disabled"}` with a 400, GPT-6 Astra rejects `none` the same way, and in Gemini 3.x the lowest level is `minimal` or `low`. You choose how much to think, not whether. In adaptive mode the model decides per request, and at lower effort it skips thinking on simple requests more often. - Thinking tokens are billed as output, even when you can’t see them. APIs usually don’t return the raw chain of thought, only a summary or an empty block with encrypted content, so `usage` shows more tokens than you can see in the response. - Thinking counts towards the output limit (`max_tokens`, `max_output_tokens`) and the context window. Set the limit too low and the response is cut off mid-thought. OpenAI recommends reserving at least 25,000 tokens for reasoning and output when you start. - The sampler is usually locked. Claude Opus from 4.7 rejects any non-default `temperature`, `top_p` or `top_k` with a 400, with or without thinking, and OpenAI with reasoning enabled doesn’t accept `temperature` or `top_p` (see “The next token”). ### What to watch out for - The chain of thought need not faithfully describe how the answer came about. In an Anthropic study (2025), models used a planted hint but admitted it only in a minority of cases: Claude 3.7 Sonnet 25% of the time, DeepSeek R1 39%. Read it as a debugging clue, not an audit. - In an agent the model also thinks between tool calls (interleaved thinking, see “The agent loop”). You send thinking blocks back unchanged in the next request, with their signature or encrypted content. The API rejects a modified block, and an omitted one breaks the continuity of reasoning. - On Claude Opus 5.5 and Fable 5.1 a block is also bound to everything before it. If your own code changes the system prompt, the tools or an earlier message (for example by trimming an old tool result), replaying the block returns a 400 on accounts created from 31 August 2026 (older ones only when they opt in). Server-side context editing and compaction don’t count as changes. Keep the history append-only and add new instructions as new messages. - Some models strip thinking from previous turns, others keep it. Kept thinking takes up the window and is billed as input. Stripped thinking changes the prefix from that point, so the cache misses after it. Changing the effort level or budget mid-conversation also invalidates the prompt cache, because the setting goes into the prompt (see “Prompt caching”). ### Check yourself **Question:** What are reasoning models, and how much thinking is worth paying for? **Short answer:** A reasoning model generates a chain of thought before answering: it lays out steps, checks them and fixes mistakes. It is the same next-token machinery, trained with reinforcement learning on tasks with verifiable outcomes such as maths or code with tests. Thinking tokens are billed as output, count towards limits and add latency. It helps on multi-step problems; on simple facts and classification it only adds cost, and it adds no knowledge. The newest models don’t let you switch thinking off, so pick the effort level with evals. The visible reasoning is not a faithful explanation of the answer. ### Follow-up questions - **How is a reasoning model different from a “think step by step” prompt?** A chain-of-thought prompt asks an ordinary model to lay out its steps, and that helps too. A reasoning model was trained to do it with RL: it writes much longer chains of thought, checks and corrects itself more often, and the API keeps the thinking separate from the answer. - **Why not always set the highest level?** Thinking tokens cost as much as output and each one adds latency, while the quality gain shrinks. On simple tasks the gain is zero, and sometimes longer thinking makes the result worse. Pick the level per task type on your own eval set. - **Does the chain of thought explain why the model answered the way it did?** There is no such guarantee. Research shows that models can use a hint without mentioning it in their chain of thought. On top of that, the API often returns only a summary. It is useful for debugging a prompt, not as an audit. - **How does thinking work in an agent with tools?** The model also thinks between tool calls, after each result. You send thinking blocks back unchanged in the next request, with their signature or encrypted content, so the reasoning stays continuous and the cache hits. On Claude Opus 5.5 and Fable 5.1, changing anything before a block ends in an error, so an agent’s history is append-only. ### Sources - [DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning](https://arxiv.org/abs/2501.12948) - [OpenAI: Reasoning models (API docs)](https://developers.openai.com/api/docs/guides/reasoning) - [Claude docs: Preserved thinking](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking) - [Anthropic: Reasoning models don’t always say what they think](https://www.anthropic.com/research/reasoning-models-dont-say-think) - [Lilian Weng: Why We Think (Lil’Log, 2025)](https://lilianweng.github.io/posts/2025-05-01-thinking/) --- ## The generation loop and KV cache *How a model writes and what it costs* *Last edited: 28 September 2026* Generation is a loop: one forward pass of the model yields one token, which is appended to the input. The KV cache holds the attention keys and values of all earlier tokens, so each step computes only the new token. The price is GPU memory, which limits context length and the number of concurrent conversations. **In plain words:** You are writing a letter, and with every new word you glance at notes on what you have already written. You don’t reread the letter, but you do look through the notes every time. The KV cache is those notes. They sit on the desk, that is, in GPU memory, and the longer the letter, the more room they take. *Interactive widget on the page: Step through with and without the cache. Compare how many tokens each step computes.* ### What the cache holds and why it works - In every attention layer a token has three projections of its state: query Q, key K and value V (see “Attention”). A new token compares its Q with the keys of all earlier tokens and takes a weighted sum of their values. The KV cache is the K and V of all tokens so far, kept separately for every layer and KV head. - Q is not kept. A token’s query is needed only in the step where that token is the last one. Later tokens read only its K and V. - Keeping them is valid thanks to the causal mask: the K and V of token *i* depend only on tokens 0…*i*, so new tokens don’t change them. This is memoisation. Without it, step *n* would recompute all *n* tokens, so generating *N* tokens would take work growing with *N*² instead of *N*. - It also works the other way round: a token’s K and V depend on all tokens before it and on its position (RoPE). Changing one token invalidates the cache for everything after it. That is why a provider can reuse a cache only for an identical prefix, and why an agent only ever appends to its history (see “Prompt caching”). ### How much memory it takes - Per token: 2 (K and V) × layers × KV heads × head dimension × bytes. Llama 3.1 70B in FP16: 2 × 80 × 8 × 128 × 2 B = 327,680 B, about 0.33 MB. A full 128k tokens is 43 GB for a single conversation, against 141 GB for the weights alone. - The architecture sets the shape of the cache. GQA keeps one K, V pair per group of query heads (MQA is the extreme case: one pair per layer). Llama 3 70B has 8 KV heads for 64 Q heads; without GQA a 128k-token cache would take 344 GB. DeepSeek’s MLA stores one compressed vector per layer: in DeepSeek-V3 that is 61 × 576 × 2 B ≈ 70 KB per token, almost 5 times less than Llama 70B, even though the model has 671B parameters. - Sliding window: some layers see only the last W tokens, so their cache never grows beyond W. Gemma 3 has five such layers (W = 1024) for every full one, which at 128k tokens comes to about 17% of the cache of the same model with only full layers. Hybrid models replace some attention layers with SSM layers that have a fixed-size state. - When serving an existing model, precision is the lever left: an FP8 cache takes half the memory. It is an approximation, so check quality on your own evals. *Interactive widget on the page: Pick a model, context length and number of conversations. See how many H100s the cache alone takes up.* ### Prefill and decode - Prefill processes the whole prompt in a single pass. It multiplies the weights by a matrix of all tokens at once, so every byte of weights read serves hundreds of tokens and the GPU computes at full speed. This is where the prompt’s cache is built. Prefill sets time to first token (TTFT), and attention grows with the square of the length, so a very long prompt increases TTFT faster than linearly. - Decode produces one token per conversation per step and multiplies the same weights by a single vector. In FP16 that is about 1 operation per byte of weights for each conversation in the batch, while an H100 needs about 300 before computation becomes the bottleneck (see “Why the GPU is idle”). Step time ≈ (weights + cache of all conversations in the batch) / memory bandwidth. - Weights are read once for the whole batch, but each conversation’s cache separately. That is why servers pack many conversations into one step (see “Continuous batching”), why one conversation’s long context slows down the whole batch, and why providers charge several times more for an output token than for an input token. *Interactive widget on the page: Increase the context and the number of conversations in the batch. See when reading the cache starts to outweigh reading the weights.* ### The cache on the server - PagedAttention (vLLM) splits the cache into blocks of a dozen or so tokens, like virtual-memory pages. The server doesn’t reserve space for the maximum length, and a prefix shared by many conversations is stored once. Prompt caching is the same blocks kept after the request and looked up by a hash of the prefix. - When memory runs out, the server pauses some conversations and later either recomputes their cache or offloads it to RAM or disk and loads it back. Both cost time but don’t change the result. - Evicting some tokens from the cache (StreamingLLM, H2O: the first, the most recent and the most attended tokens stay) saves memory but is an approximation. The model loses access to the evicted text, so quality on long tasks drops. ### Check yourself **Question:** What is the KV cache, and how does prefill differ from decode? **Short answer:** Generation is a loop: one forward pass yields one token. Attention needs the keys and values of all earlier tokens, and thanks to the causal mask they never change, so they are computed once and kept in GPU memory. Prefill processes the whole prompt in parallel, is compute-bound and sets time to first token. Decode goes token by token and is memory-bound, because every step reads the weights and the whole cache. The cache is about 0.33 MB per token for a 70B model with GQA, so it is what caps context length and batch size. GQA, MLA, sliding windows and FP8 shrink it. ### Follow-up questions - **What determines time to first token, and what determines writing speed?** Time to first token depends on the queue and on how much of the prompt has to be computed outside the cache hit, because that is prefill. Writing speed depends on the size of the weights, the context length, the number of conversations in the batch and memory bandwidth, because that is decode. - **Why does a long context slow down generation if each step computes only one token?** Because every step reads the whole cache. For a 70B model at 128k tokens that is 43 GB per step, almost a third of the weights’ size, and decode is bound by memory bandwidth. - **vLLM runs out of memory on long contexts. What do you do?** Work out the budget: weights plus cache per token × maximum context × number of concurrent sequences. Then cap the maximum length and the number of sequences, enable an FP8 cache and prefix sharing, and if that is not enough, split the model across more GPUs. Tensor parallelism splits a GQA cache by KV head, so across at most as many GPUs as there are KV heads (8 in Llama 3.1 70B). An MLA cache is replicated on every GPU, which is why DeepSeek-style models use data-parallel attention instead. - **How do you shrink the KV cache?** At model design time: GQA or MQA, MLA, sliding windows, SSM layers. At serving time: FP8, prefix sharing, offloading to RAM, evicting tokens. Only sharing and offloading leave the output unchanged. ### Sources - [kipply: Transformer Inference Arithmetic](https://kipp.ly/p/transformer-inference-arithmetic) - [How To Scale Your Model: the inference chapter](https://jax-ml.github.io/scaling-book/inference/) - [DeepSeek-V2: where MLA comes from (arXiv)](https://arxiv.org/abs/2405.04434) - [Efficient Memory Management for LLM Serving with PagedAttention (arXiv)](https://arxiv.org/abs/2309.06180) - [Sebastian Raschka: Understanding and Coding the KV Cache in LLMs from Scratch (2025)](https://magazine.sebastianraschka.com/p/coding-the-kv-cache-in-llms) --- ## Prompt caching *How a model writes and what it costs* *Last edited: 28 September 2026* Prompt caching is the KV cache kept by the provider between requests. A request that starts exactly like a recent one skips prefill for that part: it pays a fraction of the input price and gets its first token sooner. **In plain words:** A waiter who knows the regulars doesn’t ask again about your allergies and favourite table. But give a different name and they start from scratch. And they remember you for only a few minutes after your last visit. *Interactive widget on the page: Pick a prompt layout. See how much of the second request hits the cache and what it costs.* ### How it works - After a request, the provider keeps the prompt’s KV cache in blocks and indexes them by a hash of the prefix, that is, everything from the start of the prompt to the end of the block (see “The generation loop and KV cache”). A request that starts the same way loads those blocks instead of recomputing them. - Only an identical prefix hits. Each token’s K and V depend on all tokens before it and, because of the causal mask, on nothing after it. So an appended suffix leaves the cache valid, while the first changed token invalidates everything after it, even if the rest is identical. - The gain is twofold: cheaper input and a shorter time to first token, because prefill skips the cached part. It does not speed up generating the answer. The model sees exactly the same text, so the cache doesn’t change quality. ### Provider terms - Anthropic (September 2026): reads at 0.1× the input price (even less on the newest models), writes at 1.25× for a 5-minute entry and 2× for a one-hour entry. You mark the cut points yourself (up to 4 `cache_control` markers) or set a single field for automatic mode. The minimum prefix is 512 to 4096 tokens depending on the model. A shorter one is silently not cached. - OpenAI caches prefixes of 1024 tokens or more automatically. From GPT-5.6, writes cost 1.25×, reads 0.1×, and an entry lives for at least 30 minutes. Older models charge nothing extra for writes and by default keep an entry for about 30 minutes, up to 24 hours; with Zero Data Retention the default is 5–10 minutes of inactivity, up to an hour. Gemini from 2.5 has an automatic cache with no hit guarantee, and an explicit one billed per hour of storage. - Every hit renews the entry’s lifetime. At Anthropic it counts from the start of the request, so a 4-minute generation uses up most of a 5-minute entry. With such an entry, a user who replies after a quarter of an hour starts with a write. ### How to lay out the prompt - Stable parts first: tool definitions, system prompt, large documents, examples. Then the history, and the new message last. At Anthropic the hierarchy is `tools` → `system` → `messages`, so changing the tools invalidates everything. - Only ever append to the history. An agent resends all of it every turn (see “The context window and agents”), and only an append-only history lets each turn read everything but its newest part from the cache. Editing an old message, compaction, cutting an old tool result or reordering tools invalidates the cache from that point. Saving window space (see “Context engineering and memory”) therefore has to be weighed against the cost of a miss. - Typical prefix killers: a date and time or a session ID at the start, non-deterministic ordering of JSON keys or tools, switching the model, thinking level or response schema mid-conversation. - Measure hits in `usage`: `cache_read_input_tokens` at Anthropic, `cached_tokens` at OpenAI. Hit rate is cached tokens divided by all input tokens. At Anthropic `input_tokens` counts only what follows the last breakpoint, so all input is `input_tokens` + `cache_creation_input_tokens` + `cache_read_input_tokens`; at OpenAI `input_tokens` already includes `cached_tokens`. Zero with a repeated prefix means something is changing it. - Break-even: with writes at 1.25 and reads at 0.1, an entry pays for itself after a single hit (1.35 instead of 2 input prices). A one-hour entry at 2 needs two hits. A prefix that never repeats only pays the write surcharge. ### Check yourself **Question:** How does prompt caching work, and how do you structure a prompt for it? **Short answer:** It is the KV cache kept by the provider across requests. It only matches an identical prefix, because a token’s keys and values depend on everything before it: the first changed token invalidates the rest. So stable parts such as tools, the system prompt and large documents go first, history is append-only, and volatile data goes last. Reads usually cost 10% of the input price, writes can cost more than plain input, and entries live for minutes. The payoff is cheaper input and faster time to first token. The hit rate comes from the usage fields. ### Follow-up questions - **The agent bill is high and the hit rate is low. What do you check?** Diff consecutive requests token by token and look for the first difference: a date or ID at the start, non-deterministic serialisation of tools and JSON, edited history, a change of model or thinking level. Then check whether the gaps between turns exceed the entry lifetime and whether the prefix meets the minimum length. - **How is prompt caching different from a semantic cache?** Prompt caching stores the computation for an identical prefix, and the model still generates a new answer, so quality doesn’t change. A semantic cache returns an old answer to a similar question: it saves the whole call but may return the answer to a different question. - **When does caching not pay off?** When the prefix rarely repeats within the entry’s lifetime. With a write surcharge, every entry that never gets a hit costs more than a request without caching. - **Can the cache leak another user’s data?** Large providers don’t share the cache across organisations, but within your application users with the same prefix hit the same entry. A hit shows up as a shorter response time, so someone can test whether another user recently sent a given text. A 2025 audit found cache shared globally, across organisations, at 7 of 17 API providers, and at least five changed that only after disclosure. Isolate the cache per customer (a separate workspace at Anthropic; at OpenAI, from GPT-5.6, a separate prompt_cache_key, while on older models the key only steers routing) or keep sensitive data out of the shared prefix. ### Sources - [Claude docs: prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) - [OpenAI docs: prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching) - [Manus: Context Engineering for AI Agents](https://manus.im/blog/Context-Engineering-for-AI-Agents-Lessons-from-Building-Manus) - [Auditing Prompt Caching in Language Model APIs (arXiv)](https://arxiv.org/abs/2502.07776) --- ## The context window and agents *How a model writes and what it costs* *Last edited: 28 September 2026* The model remembers nothing between requests; it knows only what it received in the current one. The context window is the token limit of a single request, shared by input and output. An agent resends the whole growing history every turn, so it pays for that history many times over, and answer quality drops as it grows. **In plain words:** The model is a consultant with amnesia. Before every meeting they get a folder and know only what is in it. The folder can only be so thick, and the agent adds pages to it at every step. *Interactive widget on the page: Move the agent-turn slider. See what fills the window and how fast the cumulative input cost grows.* ### What counts towards the limit - Everything in the request: the system prompt, tool definitions (including those from MCP servers), the whole history with tool results, images and PDFs, plus everything the model generates in this turn, thinking included. It is measured in the model’s own tokens, so the same text gives different counts at different providers (see “Tokens”). You can count them before sending with a token-counting endpoint. - Part of the overhead is fixed and paid every turn: definitions for a few dozen tools easily add up to tens of thousands of tokens before the agent has done anything (see “Tools (function calling)”). ### Input limit and output limit - The window covers input and output together. A separate, much smaller output limit applies per request (`max_tokens`): as of September 2026, Claude models with a 1M-token window generate at most 128k per request, and OpenAI’s GPT-6 models have a 1.05M window, at most 922k of it input, and 128k of output. - The API usually rejects input that is too long with an error. Truncating history is a decision made by your application or framework, often silently, starting with the oldest messages. - When output hits the limit, generation stops mid-sentence: `stop_reason: "max_tokens"` in Anthropic (or `"model_context_window_exceeded"` on Claude 4.5+, when input plus output fill the window), `finish_reason: "length"` in Chat Completions. JSON is then incomplete, and a reasoning model may never get to the answer. Check the stop reason in code. - Long context can cost more per token: above 200k input tokens Gemini 3.1 Pro, and above 272k GPT-6 models, bill the whole request at 2× for input and cache and 1.5× for output. Claude models with a 1M window have a flat rate (September 2026). ### Cost grows with every turn - In the agent loop (see “The agent loop”), turn *t* sends everything from earlier turns plus the new result. With a constant increment per turn, cumulative input tokens grow with the square of the number of turns. In the widget, 24 turns already add up to about 2.2M input tokens, even though each request fits in the 200k window. - Input dominates in agents: Manus reports an average input-to-output token ratio of about 100:1. The main lever on the bill is therefore prompt caching. The causal mask keeps old tokens’ K and V fixed, so as long as the history is only appended to, every turn reads all of it but the newest part from the cache (see “Prompt caching”). Reads cost about 10% of the price, but the curve stays quadratic. - Latency grows too. Prefill of the uncached part lengthens time to first token, and every decode step reads the whole KV cache, so a long context also slows down generation (see “The generation loop and KV cache”). - While a tool runs, the model computes nothing, yet the conversation’s KV cache still occupies GPU memory. A busy server can offload the cache to CPU memory or SSD and reload it when the result arrives: faster than recomputing a long history, but not free for time to first token (Glenn Lockwood, September 2026). Through an API you see only the entry lifetime, which at Anthropic counts from the start of the request: generation plus a slow test suite can outlast a 5-minute entry, and the next turn pays for a write and a full prefill of the whole history. For tools that take minutes, the one-hour entry pays off. ### Quality drops with length - Models perform worse on long input well before the window is full. The needle-in-a-haystack test, finding one sentence that matches the question word for word, comes out almost perfect and says little. In RULER (2024), of 17 models claiming at least 32k tokens, only half kept a satisfactory score at 32k. - The drop is bigger when the question and the answer share no words: in NoLiMa (2025), 11 of 13 models fell below half of their short-context score at 32k tokens. Chroma (2025, 18 models) showed degradation from length alone even on simple tasks, stronger with similar but irrelevant passages. - Position matters too. In the Lost in the Middle study (2023), models made best use of information at the start and end of the context and markedly worse use of the middle. So put the important instruction and the question at the end, after the long material. - Window capacity is not a budget to fill. Test your use case at the lengths you actually send, and keep the window short and relevant. How to do that (trimming tool results, compaction, sub-agents, memory outside the window) is covered in “Context engineering and memory”. ### Check yourself **Question:** What counts towards the context window, and why does a long agent session get expensive and progressively worse? **Short answer:** The model is stateless and only sees the current request. The window caps tokens per request and is shared by the system prompt, tool definitions, history with tool results and everything the model generates, thinking included. The output limit is separate and much smaller. An agent resends the whole history every turn, so cumulative input tokens grow quadratically with turns. Prompt caching lowers the price, not the shape of the curve. Quality degrades with length long before the window is full, so keep context lean and test at realistic lengths. ### Follow-up questions - **What is context rot?** Quality degrading with input length before the window runs out. It is stronger when the context holds similar but irrelevant passages and when the answer doesn’t repeat words from the question, so a needle-in-a-haystack test doesn’t reveal it. - **The model has a 1M-token window. Why not put the whole knowledge base in the prompt?** Because every request pays for the whole million, at some providers at a higher rate above a threshold, time to first token grows, and quality drops with length. For a small, stable set with prompt caching it is reasonable; for a large or changing one, retrieval wins (see “RAG”). - **An agent overflows the window after 40 turns. What do you do?** First measure what fills the window, usually old tool results and tool definitions. Then trim results at the source, clear old ones, move side tasks to sub-agents, and only as a last resort compact the history (see “Context engineering and memory”). - **Does an API that keeps history server-side lower the cost?** No. Features such as previous_response_id in OpenAI’s Responses API save on transfer, but the whole history still goes to the model and counts as input. Only prompt caching or a shorter context lowers the cost. ### Sources - [Claude docs: context windows](https://platform.claude.com/docs/en/build-with-claude/context-windows) - [RULER: What’s the Real Context Size of Your Long-Context Language Models? (arXiv)](https://arxiv.org/abs/2404.06654) - [NoLiMa: Long-Context Evaluation Beyond Literal Matching (arXiv)](https://arxiv.org/abs/2502.05167) - [Lost in the Middle: How Language Models Use Long Contexts (arXiv)](https://arxiv.org/abs/2307.03172) - [Glenn Lockwood: What are KV caches really? (September 2026)](https://blog.glennklockwood.com/2026/09/what-are-kv-caches-really.html) --- ## Context engineering and memory *Context and knowledge* *Last edited: 28 September 2026* The model knows only what is in the context of the current call. Context engineering is choosing what goes in: the fewest tokens that are enough for the next step, and from memory outside the window only what is needed right now. **In plain words:** A hospital shift handover. The night shift doesn’t recount the whole night minute by minute; it leaves a chart: allergies, medication given, what has changed, what is still to do. Whatever isn’t on the chart, the day shift doesn’t know. *Interactive widget on the page: An agent has spent 40 turns adding refunds to a payments module. Choose how it manages context, and see which of six facts survive until the question in turn 41 and whether the agent answers correctly.* ### What competes for window space - A single call includes the system prompt, tool definitions, examples, retrieved documents, conversation history, tool results and the model’s reasoning (thinking). In agents, tool results grow fastest. Manus reports an average of about 100 input tokens per output token. The window limit and how cost grows with every turn are covered in “The context window and agents”. - Quality drops before the window runs out. This is context rot. Chroma (July 2025) tested 18 models: results got worse with input length even on simple tasks, and passages on the same topic that don’t answer the question did extra damage. On LongMemEval, every model did markedly better on the version with only the relevant passages (about 300 tokens) than on the full one (about 113k), even though the answer was in both. Anthropic explains this as an attention budget: n tokens make n² pairs, so attention is spread across ever more candidates (see “Attention”). Fewest tokens doesn’t mean short: a missing fact hurts as much as noise. ### The prompt: the minimum that works - Task, reason, constraints and output format. A reason works better than a prohibition. Instead of “NEVER use ellipses”, say that the text will be read by a speech synthesiser that can’t pronounce an ellipsis. From the reason the model also infers cases you didn’t list. Say what to do rather than what not to do. The test from the Claude docs: would a colleague with no context be able to follow this instruction? - Examples steer format and tone more effectively than a description. The Claude docs recommend 3–5 varied examples in `` tags. Anthropic advises a few typical ones rather than a list of every edge case. The model copies examples that are too similar together with their accidental features, such as always the same length. - Structure comes from sections or XML tags: instructions, context, examples and input data kept apart, so the model doesn’t confuse instructions with data. Long documents go at the top, the question at the end. In Anthropic’s tests this gave up to 30% better answers with multiple documents. ### Techniques for long tasks - Stable prefix first: system prompt, tool definitions, fixed documents. Variable content goes last, history is append-only, and serialisation is deterministic (the same JSON key order). Then consecutive calls hit the cache; see “Prompt caching”. Any change in the middle of the context, such as compaction, clearing or a new tool, invalidates the cache from that point. So edit rarely and in large chunks. For this reason Manus doesn’t remove tools mid-task; it masks them out during decoding instead. - Just-in-time retrieval instead of loading everything up front. The agent keeps lightweight pointers (file paths, URLs, queries) and reads the content with tools when it needs it: `grep`, `head`, a database query. Claude Code combines both: `CLAUDE.md` goes into the context at the start, and the agent reads files as it goes. Agent Skills do the same with instructions: the prefix holds only each skill’s name and description, about 100 tokens, and the full instructions load when a task matches. It is the same fix as tool search for bloated tool definitions (see “MCP”); more in “The coding-agent harness”. The price: more turns and slower than ready-made results from an index (see “RAG”). Load rules that always apply up front. Search finds what resembles the question, and “nothing on production without approval” resembles none. - Compaction: at a threshold the model summarises the history, and work continues from the summary and the latest turns. Claude Code keeps architectural decisions, unresolved bugs and implementation details, drops repeated tool results and adds the 5 most recently read files. The summary loses details that looked unimportant at the time, and producing it costs a call that reads the whole history. As of September 2026 the Claude API (beta) and OpenAI’s Responses API can compact on the server, at a token threshold or on request. Claude returns a readable summary and accepts your own summarisation prompt; OpenAI’s compaction item is encrypted, so only evals show what it lost. - Clearing old tool results is the gentlest form. A raw result from 30 turns ago is rarely needed verbatim. You replace it with a placeholder while the call itself stays, so the agent knows what it has already checked and can fetch it again. In the Claude API this is context editing: past a threshold (100k input tokens by default) it clears older results and keeps the 3 most recent. On Claude Fable 5.1 and Opus 5.5, if you send thinking blocks back, leave clearing to the API. Editing an earlier tool result in your own code, or compacting while keeping recent turns with their thinking, fails the prefix check for every later thinking block: a 400 by default for accounts created from 31 August 2026. One summary that replaces the whole history is fine. - Notes outside the window: the agent maintains a file itself, such as `NOTES.md`, a to-do list or `progress.txt`, and reads it after a context reset. Manus, averaging about 50 tool calls per task, keeps rewriting `todo.md` so the goal sits at the end of the context instead of getting lost in the middle. For coding work the Claude docs suggest sometimes starting from a clean window instead of compacting: the model rebuilds its state from notes, tests and git history. Notes contain only what the agent judged important, and they go stale too. - A sub-agent gets a clean window for a side task, such as searching a repository. It uses tens of thousands of tokens and returns a summary, according to Anthropic usually 1–2k tokens. The main context stays short, but the sub-agent doesn’t see the main thread’s decisions, so the brief it gets must include them. More in “Multiple agents”. ### Memory across sessions - Long-term memory is a store outside the model that comes back into context in later sessions. It holds facts and preferences (works in TypeScript, prefers short answers), episodes (how a similar ticket was resolved, what didn’t work) and rules (corrected instructions). It takes one of two forms: a single profile rewritten as a whole, which is simple but easily loses something on update, or a collection of small entries, which loses less but is harder to search, correct and delete. - Memory comes back into context in two ways. A small profile always sits in the prefix: simple, but you pay for it on every call. A larger store is searched by a tool when the agent needs it: it scales, but may miss. The memory tool in the Claude API runs client-side. The model requests file operations in a `/memories` directory, your code carries them out on your storage, and the API adds an instruction to always check that directory before starting work. Writing during work is visible immediately but slows things down. Writing in the background after the session doesn’t slow anything down but takes effect with a delay. - Memory goes stale. An entry saying “the project uses MySQL” after a migration to Postgres does more harm than no entry, because the model trusts its notes. Keep a date and source with each entry, let the newer one win, and let unused entries expire. The memory tool docs add a file size limit, stripping sensitive data before saving, and protection against paths like `/memories/../../`. - Memory poisoning: foreign content (a web page, document or email) tells the model to save a false entry, which then takes effect in every session. In 2024 Johann Rehberger showed false memories being planted in ChatGPT via documents, images and browsed pages, and then an entry that sent all future conversations to the attacker. According to the author, OpenAI blocked that exfiltration channel in September 2024, but not the memory writes themselves. Treat writing to memory as a privileged action: only from what the user says or with their consent, and with a visible notice. See “Prompt injection”. - Users must be able to see, correct and delete their memory, and to chat without it. One user’s memory must never reach another user’s context: a separate space per user and project, and when an account is deleted, its memory goes too. ### Check yourself **Question:** After an hour of work, an agent forgets what was agreed at the start of the session and keeps getting more expensive. How would you design context and memory management? **Short answer:** The model knows only what is in the current context, so before every call assemble the smallest set of tokens the next step needs. A stable prefix with instructions and tools goes first, for the cache. The agent fetches large data with tools when it needs it. Clear old tool results and compact history at a threshold, but have the agent write rules and key decisions to a notes file, because a summary loses details. Side tasks go to sub-agents with a clean window. Treat cross-session memory as untrusted data: dated, sourced, expiring and under user control. ### Follow-up questions - **Compaction or a fresh window?** Compaction preserves the continuity of a long conversation, but the summary loses details and costs an extra call. For code a fresh start is often better: the state lives on disk in notes, tests and git history, and the model can read it back. The Claude docs present this as an alternative to compaction. - **How would you check whether compaction loses something important?** With evals on long sessions: facts given early, questions about them after compaction, and a comparison with the answer on the full history. Tune the compaction prompt for completeness first, then cut what is unnecessary. Facts that must not be lost, the agent writes to its notes before summarisation happens. - **What goes in the stable prefix, and what do you fetch on demand?** The prefix holds rules that always apply and short facts needed in most steps. Search finds what resembles the question, and a general rule rarely resembles a specific question. Large, rarely needed data the agent fetches through paths and tools. - **A user says the assistant “remembers” something they never said. What do you check?** Where the entry came from: a conversation, a document, a web page or a tool result. If it came from foreign content, it is memory poisoning. The fix: write only from what the user says or after their confirmation, with source and date on every entry and a visible notice when something is saved. - **Why not put everything into a million-token window?** Every call pays for the whole input and waits longer for the first token, and quality drops with length and noise. In the Chroma study, every model tested did markedly better on about 300 tokens of relevant passages than on about 113k tokens in which the same answer was buried in noise. A large window is headroom, not a strategy. ### Sources - [Anthropic: Effective context engineering for AI agents (2025)](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents) - [Chroma: Context Rot (2025)](https://www.trychroma.com/research/context-rot) - [Claude docs: Memory tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool) - [Claude docs: Preserved thinking (history edits and thinking blocks)](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking) - [Embrace The Red: SpAIware, false memories in ChatGPT as an exfiltration channel (2024)](https://embracethered.com/blog/posts/2024/chatgpt-macos-app-persistent-data-exfiltration/) --- ## Embeddings and vector search *Context and knowledge* *Last edited: 28 September 2026* An embedding model turns a whole passage of text into a single vector, so that texts with similar meaning get nearby vectors. Retrieval in RAG rests on this: the question becomes a vector too, and the database returns the passages with the nearest vectors, even when they share no words with it. **In plain words:** A map where every passage of text has its own pin, and texts with similar content sit close together. Searching means pushing in a pin for the question and collecting its nearest neighbours. A contract number barely moves a pin, so contracts 48213 and 48231 sit almost on the same spot. *Interactive widget on the page: Pick a question or type your own and see which passage each method puts on top.* ### One vector for the whole passage - An embedding model, usually a transformer, computes a vector for every token and combines them into one: by averaging them (mean pooling) or, in embedders built on a decoder LLM such as Qwen3-Embedding (2025), by taking the last token’s vector. The result has a fixed length, typically a few hundred to a few thousand numbers, however long the text. This is not the same as the vectors inside an LLM from “Vectors and matrices”: there every token has its own vector, and the model learns to predict the next token, not to compare texts. - The model is trained contrastively, on pairs: a question and the passage that answers it should get nearby vectors, while the question and the other passages in the same batch should end up far apart. So “close” means “about the same thing, or answers it”, within the limits of what was in the training data. - Closeness is measured by the cosine of the angle between two vectors. If the vectors have length 1 (some models return them that way; for the rest you normalise them yourself), cosine is just the dot product, and ranking by Euclidean distance comes out identical. ### Searching millions of vectors - Exact search (brute force) compares the question with every vector. The index takes number of vectors × dimensions × bytes, e.g. 10 million passages × 1024 dimensions × 4 B (float32) ≈ 41 GB, and every query reads all of it. At tens of thousands of passages that means milliseconds and no missed results; at millions, with heavy traffic, it is too slow and too expensive. - ANN (approximate nearest neighbours) checks only a fraction of the vectors, at the cost of sometimes missing a neighbour. HNSW builds a multi-layer graph of “who is close to whom” and descends it greedily towards the question: fast and accurate, but for low latency the vectors and the graph sit in RAM. DiskANN (2019) keeps only compressed vectors in RAM and the graph with full vectors on SSD: in the paper, a billion vectors on one machine with 64 GB of RAM, under 3 ms per query. IVF splits the vectors into clusters (k-means) and searches only the few nearest ones: less memory, especially with compression, but usually lower recall at the same speed. You tune the trade-off with one parameter, `efSearch` in HNSW or `nprobe` in IVF, and measure recall against brute force on a sample of queries. - Vector quantisation cuts memory: int8 takes 4 times less, binary (1 bit per dimension) 32 times less, so about 1.3 GB instead of 41 GB. In a Hugging Face test, binary vectors alone kept about 92.5% of retrieval quality; rescoring the top candidates with the full-precision question vector against the same binary vectors raised that to about 96%, at no extra memory. The figures depend on the model. - The other lever is fewer dimensions. Models trained with the Matryoshka method keep the most important information in the first dimensions, so a stored vector can be truncated and renormalised without re-embedding the corpus. OpenAI reports that text-embedding-3-large truncated from 3072 to 256 dimensions scores better on the MTEB benchmark than the older ada-002 with 1536. ### Weak spots, hybrid search and reranking - A vector captures the general meaning and loses the details. To a vector, “contract 48213” and “contract 48231” are almost the same thing: a contract with some number. The same goes for codes, numbers, product names the model never saw in training, and company jargon. “Remote work requires approval” and “does not require approval” also get almost the same vector: in the NevIR benchmark (2023) most retrieval models, including the best ones, ranked documents that differ only by a negation no better than chance; cross-encoders did best, and a 2025 reproduction found listwise LLM rerankers better still, though below humans. - Vector search always returns k results, even when the answer isn’t in the database. A fixed similarity threshold carried over from another model won’t work, because the scale depends on the model: in multilingual-e5, scores cluster between 0.7 and 1.0. You tune the threshold on your own data, and the reranker’s score is a more reliable signal. - That is why hybrid search is the standard: BM25, the classic ranking by shared words weighted by how rare they are, plus vectors. The two lists are merged with, for example, RRF (reciprocal rank fusion): a document gets the sum of 1/(60 + rank) from each list. RRF looks only at ranks, so there is no need to reconcile BM25 and cosine scales. BM25 needs a stemmer or lemmatisation, otherwise “contract” and “contracts” count as different words; in a heavily inflected language such as Polish this is essential. - Last comes the reranker, usually a cross-encoder: it reads the question and the passage together, so it sees negation and details that two separately computed vectors can’t capture. It is too slow for the whole corpus, so it only reorders the top candidates, e.g. 50–100 of them. In Anthropic’s test (contextual retrieval, 2024), adding context to passages plus BM25 cut the share of relevant passages missing from the top 20 (1 − recall@20) from 5.7% to 2.9%, and reranking brought it to 1.9%. More in “RAG”. - Techniques that work around the weaknesses of one vector per passage (HyDE, questions generated for each passage, multi-query, small-to-big, ColBERT-style late interaction) are covered in the “Query-side and index-side tricks” section of “RAG”. ### Decisions when building the index - Chunking decides what a vector means. A long passage averages several topics and matches no question well. A short one loses context: “during this period you are entitled to…”, but which period? Prepending the document and section title before embedding helps. Text longer than the model’s limit is truncated: after 512 tokens in multilingual-e5, after 32k in Qwen3-Embedding. - A question and a document are different kinds of text: a few words versus a paragraph. Some models expect prefixes, e.g. `query: ` and `passage: ` in e5, or an input-type parameter in the API. A missing or swapped prefix throws no error; it silently degrades the ranking. - For a language other than English, such as Polish, choose a multilingual or language-specific model and test it on questions in that language, not just on the MTEB leaderboard. The PIRB benchmark (41 Polish retrieval tasks) compares more than 20 models and shows that a hybrid with keyword search improves even the best vector models. - Changing the embedding model means re-embedding the whole corpus, because vectors from two models are not comparable, even with the same number of dimensions. Store the source text and the model version with every vector, build the new index next to the old one, and switch when it wins on your test set. - Evaluate retrieval separately from answers, on a set of questions with the relevant passages labelled. Recall@k: the share of relevant passages found in the top k results. MRR: the mean of 1/rank of the first relevant result, so rank 1 gives 1 and rank 3 gives 0.33. Set k to the number of passages you actually put into the prompt. See “Evals”. ### Check yourself **Question:** How does vector search work, and why is it not enough on its own? **Short answer:** An embedding model turns a whole passage into one vector, trained so that texts with similar meaning land close together. The query is embedded with the same model, and search returns the nearest vectors by cosine similarity. At millions of passages an approximate index such as HNSW trades a little recall for milliseconds and costs memory, which quantisation reduces. Vectors catch paraphrases but miss identifiers, numbers, rare names and negation, and they always return something. So fuse them with BM25 via RRF, rerank the top candidates with a cross-encoder, and measure recall@k separately from answer quality. ### Follow-up questions - **You are changing the embedding model. What happens to the index?** You re-embed the whole corpus, because vectors from different models are not comparable, even with the same number of dimensions. You build the new index next to the old one, compare recall@k on the same question set, and only then switch traffic. That is why you store the source text and the model version with every vector. - **HNSW or IVF?** HNSW gives high recall at low latency, as long as the vectors and the graph fit in RAM. IVF splits the vectors into clusters and searches a few of them. With PQ compression it fits far more vectors into the same memory, at the cost of recall, which you win back with a larger nprobe or by rescoring the top results on full vectors. - **How do you drop passages when the answer is not in the database?** You tune a cosine threshold on your own test set, which also includes questions with no answer in the database, because the scale depends on the model. The reranker’s score is a more reliable signal. And the prompt explicitly allows the model to say “this is not in the sources”. - **How do you combine vector search with filters such as permissions or dates?** A filter applied after the search cuts results out of the top k, so with a narrow filter too few are left, or none. It is better to filter while traversing the index, which many vector databases support, or to keep separate indexes, e.g. one per customer. Enforce permissions at retrieval time, not in the prompt. - **Recall@10 is high, but the answers are still weak. Where do you look?** First, whether the relevant passage lands at ranks 8–10 while only the top three go into the prompt: a reranker helps there. Next, whether the passage, once cut out of its document, actually contains the answer (chunking). Finally, whether the model distorts it, which is a job for generation evals. ### Sources - [Malkov, Yashunin: Efficient and robust approximate nearest neighbor search using HNSW graphs](https://arxiv.org/abs/1603.09320) - [Cormack, Clarke, Büttcher: Reciprocal Rank Fusion (SIGIR 2009)](https://plg.uwaterloo.ca/~gvcormac/cormacksigir09-rrf.pdf) - [Anthropic: Introducing Contextual Retrieval (2024)](https://www.anthropic.com/news/contextual-retrieval) - [Hugging Face: Binary and Scalar Embedding Quantization](https://huggingface.co/blog/embedding-quantization) --- ## RAG *Context and knowledge* *Last edited: 28 September 2026* RAG (retrieval-augmented generation) means finding passages in your documents and pasting them into the prompt before the question. The model answers from the text it was given instead of from memory, so its knowledge can be current, private and cited, but answer quality depends mostly on what retrieval finds. **In plain words:** An open-book exam. Instead of relying on the student’s memory, you put the right page of the textbook in front of them just before the question. Give them the wrong page and they will answer wrongly with the same confidence, because they read what they were given. *Interactive widget on the page: Step through the pipeline. See which chunks drop out at retrieval and reranking, and what changes in the answer.* ### Indexing and chunking - The index is built ahead of time, off the question path: parsing documents, splitting them into chunks, embedding each chunk and storing it in a vector index and a text index. How embeddings and vector search work is covered in “Embeddings and vector search”. Parsing PDFs, tables and scans is a common source of errors: no retrieval will find text that was mangled during parsing. Layout-aware parsers and vision models help, or retrieval over page images skips text extraction altogether (ColPali, 2024). - Chunk size is a trade-off. Too small loses context (“during this period you are entitled to…”, but which period?); too large blurs the match and eats the window. A starting point is a few hundred tokens, split along the document’s structure (headings, sections, paragraphs), not by character count. You settle the final size with evals. - Contextual Retrieval (Anthropic, 2024): a model prepends 50–100 tokens of context from the whole document to each chunk before the chunk goes into both indexes. In their tests, combined with BM25 and reranking, this cut the share of relevant chunks missing from the top 20 (1 − recall@20) from 5.7% to 1.9%. - Every chunk gets metadata: source, date, version, permissions. A changed document has to be re-indexed, a deleted one removed from the index, and a change of embedding model means re-embedding the whole collection. ### Retrieval and reranking - Hybrid search (vectors plus BM25, merged with RRF) and cross-encoder reranking are covered in “Embeddings and vector search”. In the pipeline you decide the numbers: how many candidates to fetch (Anthropic took the top 150), how many chunks to put into the prompt after reranking (usually a few to a dozen or so) and how much latency you can spend on it. - A chat message is often a poor query (“so what was the deal with that leave?”). It helps to have a model rewrite it into a standalone query that takes the conversation history into account. - Permissions are filtered in the index query, based on the user’s identity and ACLs synced with the source, before a chunk reaches the prompt. An instruction like “don’t show confidential documents” in the prompt guarantees nothing, and documents can contain injected instructions (see “Prompt injection”). ### Query-side and index-side tricks - HyDE (Gao et al., 2022): a model writes a hypothetical answer, and retrieval uses its embedding instead of the question’s. It helps when questions and documents are phrased completely differently, e.g. a casual question versus the language of a policy. It costs an extra model call before every search. It hurts when the model doesn’t know the domain: made-up names and numbers pull retrieval towards the wrong neighbours. Embedding models with a separate prefix or instruction for queries partly close this gap without an extra call. - The reverse direction, on the index side: a model generates, for each chunk, the questions that chunk answers, and you index them alongside it (doc2query, 2019). You pay once, at indexing time, not on every question. The index grows, and the gain is only as large as the coverage of users’ real questions. - Multi-query extends the query rewriting from the previous section: several versions of the query, searched in parallel and merged with RRF. It raises recall when you don’t know which words the answer is written in, at the cost of a model call and several searches. A multi-step question (“who has more leave, Anna’s team or Ben’s?”) is broken into sub-questions. When each one depends on the result of the previous one, it is already agentic retrieval. - Small-to-big (parent-document retrieval): you index small chunks, because they give precise matches, and the model gets the larger whole, a section or a page, because it gives context. The price is prompt tokens, and several hits from the same section have to be merged into one passage. - Late interaction (ColBERT, 2020): a vector for every token instead of one per chunk, and relevance is the sum of the best matches of the question’s tokens against the document’s tokens (MaxSim). It sits between a fast bi-encoder and an accurate cross-encoder, at the cost of an index many times larger. ### The prompt and citations - Chunks go into the prompt with an identifier and a source, and the instruction says to answer only from them, cite the identifiers and admit when the answer isn’t there. Without that way out, the model fills the gaps from memory (see “Hallucinations”). - With long material, the documents go before the question. Anthropic reports that putting the question at the end can improve quality by up to 30% on complex inputs with many documents. The variable chunks go after a stable system prompt so they don’t break the prompt cache (see “Prompt caching”). - Check citations in code: was the cited chunk in the context, and is the quote really in it? APIs with built-in citations, such as Citations in Claude, return the cited text with a pointer to its span in the document, and the API guarantees the pointer is valid. ### Evaluation, agents and alternatives - Retrieval and generation are evaluated separately. Retrieval: a set of questions with the right chunks labelled, and recall@k, the share of the labelled relevant chunks that make the top k (with one relevant chunk: whether it is there at all), with k equal to the number of chunks in the prompt. Generation: faithfulness (does every claim follow from the chunks provided) and completeness, usually scored by a model as judge that has been checked against human ratings (see “Evals”). - Agentic retrieval: search is a tool, and the model composes the queries, reads the results and searches further on its own (see “The agent loop”). It handles multi-step questions and comparisons that a single search can’t, at the cost of several turns, time and tokens. - Long context instead of RAG makes sense for a small, stable collection, especially with prompt caching. In 2024 Anthropic put the limit at about 200k tokens (roughly 500 pages), which was then Claude’s whole window. With 1M-token windows (September 2026), cost per question, latency and quality loss with length decide well before size does (see “The context window and agents”). RAG wins for a large or changing collection, and when citations and permissions matter. Fine-tuning teaches style and format, but it teaches facts unreliably, and they are hard to update afterwards (see “Fine-tuning and LoRA”). ### Check yourself **Question:** How does RAG work, and where does it most often fail? **Short answer:** RAG gives the model knowledge through its context instead of relying on its memory. Offline, documents are parsed, chunked along their structure, tagged with metadata and permissions, and indexed both as vectors and in BM25. At query time a hybrid search runs with a permission filter, a reranker picks the few best chunks, and the prompt says to answer only from them, with citations, and to admit when the answer is missing. What fails most often is parsing and retrieval, not the model: the right chunk simply is not in the prompt. So measure retrieval recall and answer faithfulness separately. ### Follow-up questions - **Users report wrong answers. How do you find the cause?** Go through the traces stage by stage: was the right chunk in the search results, did it survive reranking, did it reach the prompt, and did the model use it? Missing from the results points to parsing, chunking or the query. Present but misused points to the prompt or the model. Stale content points to the indexing process. - **RAG or long context?** Long context is simpler for a small, stable collection, especially with prompt caching. RAG wins for a large or changing collection on cost per question, latency, permissions and citations, and a shorter context also means less quality loss from length. - **How do you handle permissions?** ACLs stored in each chunk’s metadata and synced with the source, a filter in the index query based on the user’s identity, and a test that someone without access doesn’t get the chunk. Never through an instruction in the prompt. - **When do you use agentic retrieval instead of a single search?** When the question needs several steps or a comparison of sources and can’t be answered with one query. Start with a single search, measure which questions it fails on, and pay for the agent’s turns and latency only there. - **When does HyDE hurt?** When the model doesn’t know the domain: the hypothetical answer contains made-up names, numbers and terms, so its embedding lands next to documents about something else. It also doesn’t help when searching for identifiers and exact phrases, where BM25 wins, or under a tight latency budget. Decide by comparing recall@k with and without HyDE on your own question set. ### Sources - [Anthropic: Contextual Retrieval](https://www.anthropic.com/engineering/contextual-retrieval) - [Faysse et al.: ColPali, retrieval over page images (arXiv, 2024)](https://arxiv.org/abs/2407.01449) - [Lewis et al.: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (arXiv)](https://arxiv.org/abs/2005.11401) - [Gao et al.: Precise Zero-Shot Dense Retrieval without Relevance Labels, i.e. HyDE (arXiv)](https://arxiv.org/abs/2212.10496) - [Claude docs: Citations](https://platform.claude.com/docs/en/build-with-claude/citations) --- ## The agent loop *Agents* *Last edited: 28 September 2026* An agent is a model in a loop: it gets a goal and tools, picks a call, your code runs it and sends back the result, until the model answers without a call. It differs from a workflow in one thing: the model, not your code, chooses the next step. **In plain words:** A workflow is a checklist for an intern: do this, then that, and if something doesn’t fit, come to me. An agent is an assistant who gets a goal, a phone and a calendar, and works out the next steps alone. It copes with surprises, but you don’t know in advance how long it will take or what it will do along the way. *Interactive widget on the page: Choose workflow or agent and what is in the calendar on Friday, then press ▶. Watch who picks the next step and how the context grows with every turn.* ### The loop in code - The loop is a dozen or so lines of ordinary code. Send the model the context and the tool definitions. If the response contains calls, append the model’s response to the context unchanged, run the calls, append their results and send the whole thing again. If it doesn’t, you are done. The model runs nothing; it only writes what it wants to call (see “Tools (function calling)”). - In the Anthropic API such a response has `stop_reason: "tool_use"` and `tool_use` blocks with an ID. The results go back in the next message, after the assistant’s turn, as `tool_result` blocks with the same `tool_use_id`, and you mark an error with `is_error: true`. In the OpenAI Responses API these are `function_call` and `function_call_output` items linked by `call_id`. Several calls from one response are run in parallel, and all the results are sent back together. - The model also thinks between calls (interleaved thinking; in Claude Opus 5.5 thinking can’t be turned off, as of September 2026). The Anthropic API requires the thinking blocks of a tool-use turn to come back complete and unmodified, and rejects edited ones with a 400. In the Responses API you send reasoning items back with the results, or chain requests with `previous_response_id` (see “Reasoning models”). - A response without a call is only the model’s judgement that it is done. Code adds hard stop conditions: a turn limit, a token and cost budget, a time limit, detection of repeated calls. It also handles the other stop reasons, none of which ends the task: a response cut off by the token limit (`max_tokens`) or the context window (`model_context_window_exceeded`) is an error to handle, `pause_turn` (a provider-side tool loop hit its cap) means sending the response back to continue, and `refusal` needs its own path. ### Workflow or agent - In a workflow the programmer writes the steps, and the model does narrow jobs inside them, such as extracting data from a request or writing a message. In an agent, code provides the goal and the tools, and the model picks the steps. A workflow is cheaper, faster, repeatable and easy to test, but it handles only what someone anticipated. An agent works around surprises and pays in tokens, time and predictability. The order to try: a single model call, then a workflow, and an agent only when the steps cannot be listed up front. - Patterns from Anthropic’s “Building effective agents” (2024): prompt chaining with checks in code between the calls, routing (classify, then a separate path or a cheaper model), parallelisation (independent parts at once, or several attempts and a vote), orchestrator–workers (a model splits the task into subtasks you can’t know in advance) and evaluator–optimizer (one model writes, another grades, until the result passes). In production, a workflow with one agentic step in the middle often wins. - Other names for the same ideas. ReAct (Yao et al., 2022) is this loop: the thought is the model’s text or thinking, the action a tool call, the observation its result. Andrew Ng’s four agentic patterns (2024) appear in this guide as the evaluator–optimizer (reflection), “Tools (function calling)” (tool use), the task list in long tasks (planning) and “Multiple agents” (multi-agent collaboration). ### Failures in long tasks - Return a tool error to the model as an ordinary result, with a hint: not “error 409” but “conflict with Quarterly review, check free slots”. The model then changes its approach, as in the conflict example. An exception that kills the loop wastes all the work done so far. Errors the model can’t fix, such as missing permissions or an exhausted budget, are handled by code. - In a long history the agent loses sight of its goal. What helps: a plan, i.e. a task list the agent ticks off and rewrites at the end of the context, notes in files, and history compaction (see “Context engineering and memory”). - Agents get stuck: they retry the same call, bounce between two steps, announce a success that never happened. The answer is limits and repeat detection in code, not requests in the prompt. An interrupted agent leaves the world half-changed, e.g. the meeting moved but the message not sent. So state-changing tools are idempotent, you save the loop state after every call so that work can resume after a failure, and a human gets a report of what has already happened. ### Cost, permissions and evaluation - Every turn sends the whole context again, so total input tokens grow roughly with the square of the number of turns. Latencies add up: every turn is a full model call plus the tool’s run time. The arithmetic is in “The context window and agents”. As long as the history is only appended to, everything but the newest part of each turn is a stable prefix, so with “Prompt caching” it is read from the cache at about a tenth of the input price; the curve stays quadratic, only cheaper. - An agent acts on your behalf, so it gets the least privilege it needs. Irreversible or outward-facing actions (sending, paying, deleting) are approved by a human, and code enforces that, not the prompt: in the example, code pauses the loop before the message is sent. Emails, web pages and tool results can contain instructions (see “Prompt injection”). - Judge an agent by the end state, not by its summary: is the meeting really on Friday, and did the message go out? A trace, i.e. a record of every turn with its tools, arguments, results and tokens, shows where the agent goes astray and what it costs. The same case passes one time and fails the next, so you compute pass^k: does the agent complete the task on every one of k attempts (see “Evals”). - A task that splits into independent parts can be handed out by a lead agent to sub-agents with separate contexts, at the cost of many times more tokens. See “Multiple agents”. How a coding agent puts these pieces together (permissions, sandbox, plan, compaction, subagents) is covered in “The coding-agent harness”. ### Check yourself **Question:** When would you build an agent instead of a workflow, and how do you secure its loop in production? **Short answer:** An agent is a model in a loop: it gets a goal and tools, picks a call, your code executes it and appends the call and its result, until the model answers without a call. In a workflow you write the steps and the model does narrow jobs inside them, so it is cheaper, faster and repeatable. An agent pays off only when the steps cannot be listed up front. Then code enforces turn, budget and time limits, grants least privilege, holds irreversible actions for human approval and checkpoints state after every step. Quality is measured by the end state, with each task run several times. ### Follow-up questions - **How does an agent know it is done?** It doesn’t know, it judges: it answers without calling a tool. Sometimes it announces a success that never happened. That is why code has hard limits (turns, tokens, time), and you check the result with an end-state test, not the model’s summary. - **The agent keeps calling the same tool. What do you do?** As a quick fix, code detects a repeated call with the same arguments, adds a note for the model, and breaks the loop on the next repeat. The cause is usually in the tool: an unclear error message or a result that shows no progress. The run goes into the eval set as a new case. - **The agent stopped halfway through a task. How do you design for that?** You assume it will happen. State-changing tools accept an idempotency key, and you save the loop state after every call, so work can resume without repeating side effects. A human gets a list of what has already been changed and, where possible, an action that undoes it. - **How do you judge whether an agent is ready for production?** A set of tasks from real cases, each run several times and graded on the end state, plus cost and time per task. You compute pass^k, because users expect it to work every time: for a task the agent completes in 90% of attempts, with independent attempts, three successes in a row happen in about 73% of cases. Across a suite you compute pass^k per task and average it. - **How do you classify an agent’s failures?** Following Chip Huyen (2025), into three groups. Planning: a wrong or non-existent tool, bad arguments, a missed goal, a false claim of success. Tool: the call is right, but the tool returned a wrong result, so each tool is tested on its own. Efficiency: the task is done, but with too many steps, too much cost or too much time against a baseline. Traces tagged this way show whether to fix the prompt and tool descriptions, the tool itself or the limits. - **Do you need a framework to build an agent?** No. The loop is a dozen or so lines on a plain API, and that is the place to start. A framework helps with durable state, resuming and traces, but it hides the prompt and the context, which you have to understand anyway. ### Sources - [Anthropic: Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) - [Claude docs: How tool use works (the loop and stop reasons)](https://platform.claude.com/docs/en/agents-and-tools/tool-use/how-tool-use-works) - [Anthropic: Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) - [Chip Huyen: Agents (2025, incl. failure modes)](https://huyenchip.com/2025/01/07/agents.html) - [Hugging Face: Agents Course (hands-on)](https://huggingface.co/learn/agents-course/unit0/introduction) --- ## Tools (function calling) *Agents* *Last edited: 28 September 2026* The model runs no code. It gets a list of tools, each with a name, a description and an argument schema, and when it decides it needs one, it writes a call instead of an answer: your code runs it and sends back the result. **In plain words:** A manager who can only make phone calls. They have a card listing the jobs an assistant can handle, with a short description of each. They say “check whether tomorrow’s 8:15 train is running on time”, and the assistant checks and calls back with the result. When the descriptions on the card are vague, the manager asks for the wrong thing. *Interactive widget on the page: Change the quality of the tool descriptions and toggle parallel calls, then step through the exchange. Watch which tool the model reaches for, how many requests it costs and whether the answer is true.* ### What a call looks like - The model has no access to the internet, a database or your computer. All it can do is write, in an agreed format, “call weather_forecast for Manchester for tomorrow”. The result comes back to it as another message, and it reads it like any other text. How this turns into multi-step work is covered in “The agent loop”. - Definitions go into the prompt as text in a format the model knows from training, like the roles in “How a model sees a chat”. A call is ordinary tokens too: the API extracts it and returns it as a separate field with an ID, a name and arguments, and the stop reason says the model is waiting for a result. You run several calls from one response concurrently and send back all the results, each with its call’s ID. When order matters, you turn parallel calls off (`parallel_tool_calls: false` in OpenAI, `disable_parallel_tool_use` in Anthropic). - `tool_choice`: `auto` (the default, the model decides), `required` or `any` (it must call one of them), a specific tool (e.g. for extracting data into a schema), `none` (calls not allowed). Forcing a call on every turn won’t let the model finish, so in a loop it is used only at selected steps. Forcing doesn’t work everywhere (as of September 2026): Claude Opus 5.5 and Fable 5.1 reject `any` and a specific tool with a 400 error. To extract data into a schema there, use structured outputs (see “Enforcing output format”). - Some tools are run by the provider: web search or running code in a sandbox happens on their side, and you get the finished result. Less code to write, less control over what ran and where. - Computer use and browser agents run the same loop with a screen as the tool: the model gets a screenshot, writes an action (click at x, y, type text), your code performs it and returns a fresh screenshot. Each screenshot is image input, about 1,000–1,800 tokens at Anthropic, and stays in the history, so long sessions prune old ones. Every page the agent looks at is untrusted input: text on screen can steer it like any tool result (see “Prompt injection”). ### The description is the only manual - The model chooses a tool only by its name, description and schema; it never sees the code. Write the description as you would for a new team member: what the tool does, when to use it, when not to, what format the arguments take and what it returns. For the model, two tools with similar descriptions are a coin toss, as the overlapping-tools variant shows. - What helps: names prefixed with the service (`rail_status`, `rail_timetable`), unambiguous parameters (`user_id` rather than `user`), enums instead of free text, and example calls: in Anthropic’s tests, examples in the definition raised accuracy on complex arguments from 72% to 90%. A few tools built for specific tasks beat a wrapper around every API endpoint. - You design the result too. Return only the fields that are needed, readable names instead of internal UUIDs, and trim long lists with a note on how to fetch the rest. Every token of a result stays in the context and is sent on every later turn, unless your harness clears or compacts it (see “Context engineering and memory”). - Strict mode (`strict`) is the constrained decoding from “Enforcing output format”: the arguments always match the schema. It doesn’t guarantee sensible values: the model can still guess a missing parameter or pick the wrong day. ### Every tool costs on every request - Definitions are text attached to every request, so you pay for them on every turn, even when no tool gets used. The provider adds its own hidden system prompt for tool use on top: a few hundred tokens at Anthropic, depending on the model. - Definitions sit at the start of the prompt, so they cache well (see “Prompt caching”). Adding, removing or reordering a tool mid-conversation invalidates the cache for everything after them. OpenAI has `allowed_tools` for this: it narrows the choice on a given turn without changing the list of definitions. - More tools mean a higher bill, less room in the context and a harder choice. OpenAI recommends fewer than 20 functions at a time, noting that this is a soft suggestion. Anthropic gives an example of five MCP servers with 58 tools whose definitions took about 55k tokens before the first question was asked (see “MCP”). At that scale tool search helps: the model sees a tool for searching tools and gets the full definitions only when it needs them. ### Errors and security - An error is a result too. Instead of aborting, send the model a specific message: what is wrong and how to fix it, and the model usually corrects itself on the next step. Check the arguments in code before running anything. Protect state-changing operations with an idempotency key: after a network error, your code or the model may retry the same call, and the ticket must not be bought twice. - A tool is a permission. The model can call it with wrong arguments or under the influence of text it has just read, because tool results (web pages, emails, documents) are untrusted data. Grant the least privilege needed, and make irreversible actions, such as a transfer, a deletion or a send, wait for human confirmation, enforced in code, not in the prompt. More in “Prompt injection”. ### Check yourself **Question:** How does function calling work, and how do you design tools for a model? **Short answer:** The model executes nothing. With the prompt it gets tool definitions: a name, a description and a JSON Schema for the arguments. When it needs a tool, it emits a call in a trained format instead of an answer, sometimes several at once. Your code validates the arguments, runs the call and returns the result with the call ID. Selection depends mostly on the descriptions, so write them like documentation, avoid overlapping tools and keep results concise. Definitions cost tokens on every request. Errors go back as results with a hint, state-changing operations are idempotent, and irreversible ones need human approval. ### Follow-up questions - **How does function calling differ from structured output?** From the model’s side it is the same mechanism: text in a prescribed format. The difference is in what happens next. Structured output is the final answer for your code. A tool call is a request, after which the result goes back to the model and the conversation continues. - **How many tools is too many?** There is no hard limit. The signals are wrong tool choices on your eval set and the growing cost of definitions. Then you merge similar tools, split the work between sub-agents with their own tool sets, or load definitions on demand through tool search. - **What do you do when the model passes wrong arguments?** First, improve the description and the schema: format, examples, enums. Then validate in code and send a readable error back to the model so it can correct itself. Strict mode removes syntax and type errors, but not wrong values. - **A tool result is 50k tokens. What do you do?** You don’t paste it in whole, because it would stay in the context and be sent on every turn until something clears it. You filter the fields, paginate, return a summary with a handle to the full data, and write large content to a file the agent can search with a separate tool. - **When do you give the model code execution instead of many calls?** When the task means many calls plus data processing, e.g. fetch 50 records, filter them, count them. The model writes a script that calls the tools itself, and only the final result comes back into the context. Anthropic reports about 37% fewer tokens on complex research tasks. The price is a sandbox for running code and harder auditing. ### Sources - [Anthropic: Writing effective tools for agents](https://www.anthropic.com/engineering/writing-tools-for-agents) - [OpenAI: Function calling (docs)](https://developers.openai.com/api/docs/guides/function-calling) - [Anthropic: Advanced tool use (tool search)](https://www.anthropic.com/engineering/advanced-tool-use) - [Claude docs: Define tools (tool_choice and limits on forcing)](https://platform.claude.com/docs/en/agents-and-tools/tool-use/define-tools) - [Claude docs: Computer use tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool) --- ## MCP *Agents* *Last edited: 28 September 2026* MCP (Model Context Protocol) is an open standard for connecting tools and data to model-powered applications. You write an integration once, as an MCP server, and use it in Claude, ChatGPT, an IDE or your own agent. The model knows nothing about MCP: it still sees tool definitions in the prompt and writes calls. **In plain words:** A wall socket. The kettle maker doesn’t know your wiring, and the electrician doesn’t know the kettle: a shared plug shape is enough. But the socket doesn’t check what you plug into it, so connecting faulty equipment is on you. *Interactive widget on the page: Connect MCP servers and watch what lands in the model’s context before the user types anything.* ### Why a standard: M + N instead of M × N - Without a standard, every model-powered application writes its own integration with every system: M applications and N systems mean M × N integrations. With MCP, each system gets one server and each application one client, so M + N is enough. Anthropic published MCP in November 2024 and in December 2025 handed it over to the Agentic AI Foundation under the Linux Foundation, so the standard doesn’t belong to a single vendor. - MCP standardises the application side: how to discover a server’s tools (`tools/list`) and how to call them (`tools/call`). The model side doesn’t change. The host turns the server’s definitions into ordinary tools in the model API, the model writes a call, and the host forwards it to the right server and hands back the result. The mechanism itself is covered in “Tools (function calling)”. - MCP connects an agent to tools and data; A2A (Agent2Agent, from Google, a Linux Foundation project since June 2025) connects agents to each other. An A2A agent publishes an Agent Card with its skills, and another agent hands it a task and gets back status updates and artefacts, without seeing its tools or prompt. As of September 2026 A2A has over 150 supporting organisations and integrations in the Google, Microsoft and AWS agent platforms, but no public usage figures. Subagents inside one application don’t need it (see “Multiple agents”); it pays off when the other agent belongs to another team or vendor. ### Host, client, server - The host is the model-powered application, e.g. Claude Desktop, an IDE or your agent. It creates a separate client for each server, and the host decides what reaches the model. A server doesn’t see the conversation or the other servers; it only gets the calls addressed to it. Messages are JSON-RPC 2.0 over one of two transports: stdio, when the host launches the server as a local process and talks to it over stdin and stdout, or Streamable HTTP for remote servers, where every message is a POST to a single endpoint. - A server exposes three kinds of things, which differ in who decides to use them. Tools are chosen by the model. Resources, e.g. a file or a database schema, are attached to the context by the application. Prompts, i.e. ready-made templates, are chosen by the user, usually as a slash command. In the other direction, a server can ask the user for input or consent during a call (elicitation): with a form, or, for passwords and payments, with a link opened outside the application. Sampling (the server asks the host’s model for a response), roots (the host points the server to directories) and protocol-level logging are deprecated as of version 2026-07-28. They work for at least 12 more months; the replacements are tool parameters, direct calls to the provider’s API, and stderr or OpenTelemetry. Optional extensions, negotiated through capabilities, add Tasks (a long-running call returns a handle the client polls) and MCP Apps (interactive UI the host renders in the conversation). - Specification version 2026-07-28, current as of September 2026, made the protocol stateless. There is no `initialize` handshake and no session any more: every request carries the protocol version and the client’s capabilities in `_meta`, and `server/discover`, which every server must implement and a client may call, returns the server’s versions and capabilities. As a result, a remote server scales behind an ordinary load balancer. When the server needs something from the user, it doesn’t send a request of its own; it replies with `input_required`, and the client retries the call with the answer. Older servers (up to version 2025-11-25) start with `initialize` and capability negotiation and keep a session. A client that wants to talk to them detects the version and switches to the old mode. The reverse doesn’t work: a client on 2025-11-25 or earlier has no way to switch forward, so a server that must also serve such clients has to support both eras and keep answering `initialize`. - A remote server authorises requests with OAuth 2.1 and acts as the resource server. To a request without a token it responds with 401 and the address of its metadata. In the metadata the client finds the authorisation server, takes the user through sign-in and gets a token issued for this one server. The server must not accept a token issued for another service, nor pass the client’s token on to an API (token passthrough). A local stdio server takes its credentials from environment variables. *Interactive widget on the page: Step through an exchange with a calendar server. Watch which part is the MCP protocol and which is ordinary function calling.* ### Every server costs tokens - A simple host attaches the definitions of all tools from all connected servers to every request. Anthropic describes five servers with 58 tools that took about 55k tokens before the first question was asked. You pay for this on every turn, and the model makes more mistakes when tools have similar names, e.g. `notification-send-user` and `notification-send-channel`. - Limit this on the host side: connect only the servers you need and enable only the tools you need from them. With a large catalogue, the host defers the definitions: the model gets a tool for searching tools, and full definitions enter the context on demand. As of September 2026 Claude Code, Codex and Cursor do this by default; Claude Code loads only tool names and server instructions at the start. The MCP docs recommend this mode once definitions take up 1–5% of the window. Don’t reorder or remove tools mid-conversation, because that invalidates the prompt cache (see “Prompt caching”); definitions found by search are appended after the cached prefix, so they don’t break it. - Design the server around tasks, not API endpoints: one `schedule_event` that finds a free slot and creates the meeting, instead of three wrappers, `list_users`, `list_events` and `create_event` (Anthropic’s example). Return concise results with only the fields the model needs, filtered and paginated on the server side, because a result stays in the context for every later turn. When the host defers definitions, the tool name, description and server instructions decide whether the model finds your tool at all, so put the key facts first. When an agent chains many calls, code mode helps: the model writes a script that calls the tools in a sandbox, and only the final result comes back into the context. ### Someone else’s server, code and text - A local server runs with your user’s permissions and can read files or SSH keys. A remote one receives every argument the model passes to it. Tool descriptions and tool results go into the prompt as text, and the model can’t tell them apart from your instructions (see “Prompt injection”). - Tool poisoning is an instruction hidden in a tool description. The model reads it even if it never calls that tool: in every request when the host loads definitions up front, or once a search returns it, while the user usually sees only the name and a shortened description. In April 2025 Invariant Labs showed an `add` tool whose description got an agent in Cursor to read the SSH key and the MCP configuration and send them in an extra parameter. A malicious description can also change how the tools of other, trusted servers are used (shadowing). - Rug pull: a server changes its definitions after you have approved it. It only takes the next `tools/list` returning a different description, or a new package version doing something other than the previous one. - A malicious server supplies two parts of the lethal trifecta by itself: untrusted text and an exfiltration channel, because everything the model writes into its tools’ arguments goes to the server’s owner. Private data comes from any other connected server. Overly broad permissions multiply the damage: a token for all repositories turns one successful attack into a leak of everything. - The defence lives in the host, not in the prompt: an allowlist of servers, only the tools you need from them, pinned versions and fresh approval when definitions change, local servers in a sandbox, tokens with minimal scope widened on demand. Call tools that change or send something only after approval by a human who sees the arguments. Annotations such as `readOnlyHint` are declared by the server itself, so from an untrusted server they mean nothing. ### When you don’t need MCP - One application with a few internal tools: plain function calling in the application code is enough. MCP adds a separate process or service, a transport, authorisation and versioning, and the M + N gain only appears once several applications use the same integration. - MCP pays off when an integration has to work in many hosts (your products, your team’s IDEs, your customers’ Claude and ChatGPT) or when you want to use a vendor’s ready-made server instead of writing the integration yourself. The OpenAI API (the `mcp` tool in the Responses API) and the Anthropic API (the MCP connector) connect to a remote server themselves, so you don’t have to write a client. ### Check yourself **Question:** What is MCP, and when would you use it instead of plain function calling? **Short answer:** MCP is an open JSON-RPC protocol that standardises the application side of function calling: the host discovers a server’s tools with tools/list and invokes them with tools/call. The model notices nothing; it still sees tool definitions in the prompt and emits calls. The gain is M + N integrations instead of M × N: a server is written once and works in Claude, ChatGPT or an IDE. The cost is the connected servers’ definitions in the context (all of them on every request, unless the host defers them through tool search), and a new trust boundary, because a third-party server is foreign code and foreign text in the prompt. For one app with a few internal tools, plain function calling is enough. ### Follow-up questions - **You have 30 MCP servers and the agent starts picking the wrong tools. What do you do?** Measure how many tokens the definitions take and which tools overlap. Keep only the tools the task needs, and load the rest through tool search or split them between sub-agents with their own sets. Don’t reorder or remove tools mid-conversation, so as not to break the prompt cache. - **How do you allow MCP servers in a company without opening a path to data leaks?** An allowlist of approved servers with pinned versions and a diff of the definitions on every update, local servers in a sandbox, and minimally scoped OAuth tokens issued for a specific server. Write and send actions are approved by a human in the host, and every call goes to an audit log. - **How do tools, resources and prompts differ?** In who decides to use them. The model chooses tools, the application attaches resources, e.g. a file or a database schema as context, and the user picks prompts, usually as a slash command. Not every host supports all three: the MCP connector in the Anthropic API supports only tools. - **Why did version 2026-07-28 remove sessions and initialize?** So that a remote server scales like an ordinary API. Every request carries the protocol version and the client’s capabilities, so it can hit any instance behind a load balancer with no shared state. The server passes state between calls explicitly: a tool returns a handle, and the model passes it in the next call. - **You are building an MCP server on top of an existing REST API. How do you approach it?** Don’t map endpoints one to one. Pick a few tasks the agent actually performs and turn them into tools that combine several API calls themselves. Return only the fields needed, filter and paginate on the server side, and send validation errors back as a result with isError and a hint on how to fix the arguments. ### Sources - [MCP specification, version 2026-07-28](https://modelcontextprotocol.io/specification/2026-07-28) - [MCP: Client Best Practices (tool search, code mode)](https://modelcontextprotocol.io/docs/2026-07-28/develop/clients/client-best-practices) - [MCP: Security Best Practices](https://modelcontextprotocol.io/docs/2026-07-28/tutorials/security/security_best_practices) - [Invariant Labs: Tool Poisoning Attacks (2025)](https://invariantlabs.ai/blog/mcp-security-notification-tool-poisoning-attacks) - [Linux Foundation: A2A after its first year (April 2026)](https://www.linuxfoundation.org/press/a2a-protocol-surpasses-150-organizations-lands-in-major-cloud-platforms-and-sees-enterprise-production-use-in-first-year) --- ## Enforcing output format *Agents* *Last edited: 28 September 2026* Structured output in strict mode doesn’t ask the model for a format; it enforces it: at every step the sampler zeroes the probability of tokens that would break the schema. You get a guarantee of syntax, not of correct content. **In plain words:** A form with checkboxes instead of a blank sheet. You can’t write anything outside the allowed options, but you can still tick the wrong box. And when there is no box for the right answer, you tick the nearest one. *Interactive widget on the page: Step through generation. See which tokens the mask cuts, and stop at token 4.* ### How masking works - The schema is compiled into an automaton. For regular expressions and non-recursive schemas a finite automaton is enough; recursive schemas and arbitrary grammars need a pushdown automaton. The automaton’s state says which characters may come next. - Tokens don’t line up with JSON syntax: a single token might be `{"` or `":"`. So the engine (Outlines, XGrammar) works out, for each automaton state, which tokens in the vocabulary are allowed. The rest get a logit of minus infinity, and the sampler draws from what remains (see “The next token”). - The per-token overhead is close to zero, because most checks are precomputed. You pay on the first use of a schema: compilation adds latency, after which the grammar is cached (at Anthropic, for 24 hours since its last use). ### Modes in the API - JSON mode guarantees valid JSON, but not conformance to a schema. Strict mode guarantees the schema: `json_schema` with `strict: true` in OpenAI, `output_config.format` in Anthropic. - Tool calling is the same JSON in a different role: the model writes a tool name and arguments (see “Tools (function calling)”). Without strict mode the format is only learned, so the arguments usually, but not always, match the schema. `strict: true` on a tool definition turns on the same masking. In OpenAI’s Responses API tools are strict by default when the schema allows it and otherwise silently fall back to best effort (the response shows `strict: false`), so set the flag explicitly. - Strict modes support a subset of JSON Schema. OpenAI requires every field to be listed in `required` (an optional field is a type that allows `null`) and every object to have `additionalProperties: false`, with limits of 5,000 properties and 10 levels of nesting. Anthropic doesn’t support, among other things, recursive schemas or string length limits, and allows at most 20 strict tools per request, 24 optional parameters and 16 union-typed parameters across all strict schemas. Beyond that the API returns “Schema is too complex for compilation”. - The mask covers only the answer. A reasoning model’s thinking stays unconstrained, so the model can reason freely first and then fill in the schema (see “Reasoning models”). ### What the guarantee doesn’t cover - Content. The model can still write a wrong invoice number or a made-up date in a valid format. Business rules (ranges, totals, whether IDs exist) are checked by code, and the error goes back to the model with a specific message, with a cap on retries. - A response cut off by the length limit (`stop_reason: "max_tokens"`, `finish_reason: "length"`) or a refusal: OpenAI returns it in a separate `refusal` field, Anthropic as `stop_reason: "refusal"`, and the content may not match the schema. Check the stop reason before parsing. - Letter case in enums at Anthropic. String `enum` and `const` values can come back with different capitalisation (“Conversation Topic 3” for “Conversation topic 3”), with a normal stop reason and no error, in JSON outputs and strict tool use alike. Compare them case-insensitively and avoid values that differ only in case. - What the model wanted to say. When the most likely answer lies outside the schema, the mask pushes the model into the nearest allowed one, like phishing turned into spam in the widget. Add an escape hatch: an “other” or “unknown” value and a comment field. - Reasoning quality. The model writes left to right, so putting the decision before the justification makes it decide before it has worked anything out. Put the justification field before the decision field and list both in `required`: Anthropic writes required properties before optional ones (OpenAI keeps the schema’s order and requires every field anyway). In the “Let Me Speak Freely?” study (2024), format restrictions lowered scores on reasoning tasks, partly for this very reason: on one task, GPT-3.5 in JSON mode put the answer before the reason every time. The result is contested. A re-run with matched prompts (.txt, “Say What You Mean”) found no drop, while other work measures a cost from the mask itself, mainly with small models and tight schemas (Reddy et al., 2026). When your evals show a drop, let the model answer freely and extract the structure with a second, cheap call. ### Check yourself **Question:** How do you guarantee that a model returns JSON matching a schema, and what does that guarantee not cover? **Short answer:** In strict mode the schema is compiled into a grammar, and at every step the sampler zeroes out tokens that would break it. The JSON matches the schema unless the length limit cuts it off or the model refuses, which the stop reason shows. At Anthropic, enum values can also come back in a different letter case, so compare them case-insensitively. The guarantee covers syntax, not content: values are validated in code. The schema gets an escape hatch, because a tight enum pushes the model into the nearest option, and the reasoning field goes before the decision, because the model writes left to right. Strict tool calling works the same way. ### Follow-up questions - **Can enforcing a format lower quality?** Yes, when the schema asks for the decision before the reasoning or has no “other” option. What helps: a reasoning field before the decision, a reasoning model’s thinking, which the mask doesn’t cover, or two steps: a free-form answer first, then extraction. - **JSON mode, structured output or tool calling: when do you use which?** Structured output when the answer feeds your code. Tool calling when the model has to choose an action from several, with strict mode for the arguments. JSON mode only where there is no strict mode. - **How do you do this on your own model?** vLLM and SGLang have built-in constrained decoding engines (including XGrammar and llguidance): the schema turns into an automaton, and the automaton into a token mask at every step. The per-token overhead is small; the cost is compiling each new schema. With a reasoning model, start the server with a reasoning parser (e.g. --reasoning-parser in vLLM); otherwise the grammar applies from the first token and cuts off the thinking. - **The JSON parses, but the data is wrong. What next?** Validation in code (constrained types, business rules) and a retry with a specific error message, with a cap on attempts. A recurring error is a signal to change the schema or the prompt, and the case goes into the eval set. ### Sources - [OpenAI docs: Structured Outputs](https://developers.openai.com/api/docs/guides/structured-outputs) - [Claude docs: Structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs) - [Willard, Louf: Efficient Guided Generation for Large Language Models (arXiv)](https://arxiv.org/abs/2307.09702) - [Tam et al.: Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models (arXiv)](https://arxiv.org/abs/2408.02442) - [.txt: Say What You Mean, a response to “Let Me Speak Freely?”](https://blog.dottxt.ai/say-what-you-mean.html) --- ## The coding-agent harness *Agents* *Last edited: 28 September 2026* A coding agent is a model plus a harness: everything around the model that turns it into a working agent. The model proposes the next step. The harness decides what the model sees, what actually runs and what can be undone. **In plain words:** A skilled contractor on a building site. The skill is theirs, but the site decides the rest: which tools are laid out, what the job sheet and site rules say, which rooms are locked, who has to sign off before a wall comes down, and whether there is a spirit level to check the work. The same contractor does very different work on a well-run site and on a chaotic one. *Interactive widget on the page: Replay one task in a coding agent: partial refunds are 1 cent short. Switch off one part of the harness, press ▶ and see what changes.* ### What the harness is made of - The loop from “The agent loop” with a few general tools: read a file, edit by replacing a string, run a shell command, search with grep and glob. The harness validates every call, runs it and clips long output: Claude Code writes an MCP result above 25k tokens to a file and gives the model the path (Claude Code and Codex details on this page are as of September 2026). - Instructions. The system prompt describes the tools and working rules, and a project memory file is loaded at the start of every session. AGENTS.md is an open format read by Codex, Cursor, Copilot, Jules and others. Claude Code reads CLAUDE.md, or AGENTS.md when there is no CLAUDE.md. The file is context, not configuration: the model usually follows it, but nothing forces it to. - Planning. The model writes a todo list and ticks it off, and the harness puts the list back at the end of the context whenever it changes, so the goal doesn’t sink into the middle of a long history. In plan mode the agent may only read and run read-only commands until you accept the plan. - Context management, covered in “Context engineering and memory”: clearing old tool results, compaction at a threshold, and subagents that search in their own window and return a summary. Two mechanisms keep material out of the window until it is needed. Tool search puts only tool names in the prefix and loads full definitions on demand, the Claude Code default. Agent Skills, an open format, are folders with a SKILL.md file: only the name and description sit in the prefix, and the body loads when the task matches (progressive disclosure). - Permissions decide which calls wait for a human. In Claude Code, reads and read-only commands such as `ls`, `grep` and `git status` run freely, while other shell commands and edits ask by default. Modes range from plan, through accept-edits and auto (a classifier reviews each action), to bypass. A sandbox is enforced by the operating system (Seatbelt on macOS, bubblewrap on Linux): writes only in the working directory, network only through a proxy with a domain allowlist. Codex in its default workspace-write mode has the network off. Both layers are needed: without network isolation a hijacked agent can send out your SSH keys, and without filesystem isolation it can plant something that opens the network later. - Hooks are your commands at fixed points in the loop: before a tool call (they can block it), after an edit (a formatter), when the agent wants to finish (tests). They run every time, unlike an instruction the model may skip. In Claude Code a hook that denies a call blocks it even in bypass mode. MCP servers add outside tools (see “MCP”). - Checkpoints. Claude Code snapshots files before each of your prompts, and `/rewind` restores code, conversation or both. It doesn’t track files changed by shell commands such as `rm` or `mv`, nor effects outside your machine, so git stays the real undo. ### Same model, different harness - SWE-bench and Terminal-Bench measure a model and a harness together. Anthropic noted in January 2025 that scores vary significantly with the scaffold even for the same model. LangChain (February 2026, a report on its own product) kept GPT-5.2-Codex fixed and changed only the system prompt, tools and hooks, including a checklist that forces verification before the agent finishes and loop detection: Terminal-Bench 2.0 rose from 52.8% to 66.5%. It works the other way too: in spring 2026 three Claude Code changes made the same models worse for weeks (see “Compute and ‘getting dumber’”). More machinery isn’t automatically better: mini-SWE-agent, about 100 lines of Python with bash as its only tool, scores over 74% on SWE-bench Verified according to its authors. A leaderboard number describes a model–harness pair, so compare models in your own harness on your own tasks. - The harness also sets most of the bill. Every turn resends the whole context, and in agents input tokens outnumber output roughly 100 to 1. So a harness orders each request from the most stable part to the most volatile: system prompt with tool definitions, then project memory, then the conversation, which only grows at the end. A cache read costs about 10% of the input price, and one change near the start of the prefix makes the whole history full price again (see “Prompt caching”). That is why Claude Code appends plan mode and skills as messages and applies a CLAUDE.md edit only after `/clear`, `/compact` or a restart, while switching models or compacting rebuilds the cache. ### Working with a coding agent - Verification is the ground truth. Give the agent tests, a type checker, a linter and a build it can run itself, ideally a fast target such as `make test-fast`. The agent’s “done” is a claim; the exit code is a fact. - Keep AGENTS.md small and precise: build, test and lint commands, conventions the code doesn’t reveal, and what not to touch. The Claude Code docs suggest under 200 lines, because a longer file costs context and is followed less reliably. Rules that must hold go into permissions and hooks, not prose. - Plan before editing and scope the task: one bug or feature per session, with a clear criterion for done, and a fresh context between unrelated tasks. - Review the diff, not the summary. Commit before a large change, so every agent edit can be reverted with one command. - Match the permission level to the risk: read-only or plan mode in an unfamiliar repository, accepted edits inside a sandbox for daily work, and full autonomy (bypass, danger-full-access) only in a container or VM without secrets or production credentials. ### Failure modes - Losing the goal: after dozens of turns the task from the first message sits in the middle of the context or in a summary. A todo list, notes in files and shorter sessions help. - Edits without tests: in LangChain’s traces the most common failure was an agent that wrote a solution, re-read it and stopped without running anything, and Anthropic (November 2025) saw agents mark features as finished without testing them end to end. A hook that runs the tests before the agent may finish helps. - Context rot: quality drops as the history fills with logs and old file reads, long before the window is full (see “The context window and agents”). Compact at natural breaks between tasks and send searches to subagents. - Prompt injection: a README, a code comment, an issue, a dependency’s docs or a web page is text in the same context as your instructions (see “Prompt injection”). In May 2025 Invariant Labs showed an issue in a public repository that made Claude 4 Opus, connected to the GitHub MCP server, read a private repository and publish its contents in a public pull request. The defence is in the harness: approval for actions, a network allowlist and tokens with minimal scope. - Runaway cost: an agent retrying the same fix, subagents fanning out, dozens of tools loaded up front, a prefix that keeps changing. Set turn and budget limits in the harness and watch the ratio of cache reads to cache writes. ### Check yourself **Question:** What does a harness add to the model in a coding agent, and why does the same model score differently in two harnesses? **Short answer:** A harness is everything around the model that makes it an agent: the loop, a few general tools (read, edit, shell, search), the system prompt and a project file such as AGENTS.md, a todo list, context management (clearing old results, compaction, subagents, tool search), skills loaded on demand, permissions and a sandbox, hooks and checkpoints. The model only proposes calls; the harness decides what it sees, what runs and what can be undone. So tools, prompts and a verification loop move scores: LangChain reports 13.7 more points on Terminal-Bench 2.0 after changing only the harness. The harness also keeps the prefix stable for prompt caching, which sets most of the bill. ### Follow-up questions - **Why doesn’t “never run git push” in AGENTS.md protect you?** The file is context, not configuration. The model usually follows it, but a long session, an ambiguous request or injected text can outweigh it. A rule that must hold goes where the harness enforces it: a deny rule, a hook that blocks the call before it runs (in Claude Code this works even in bypass mode), a sandbox, or a token without push rights. - **The sandbox is on. How can the agent still destroy your work?** The sandbox limits writes to the working directory and the network to allowed domains, and the repository is inside that boundary, so git clean -fdx or rm -rf src goes through. Claude Code checkpoints don’t track changes made by shell commands. What helps: frequent commits, an ask or deny rule for destructive commands, and a separate worktree or container for risky tasks. - **The agent keeps reporting success while CI fails. What do you change in the harness?** A fast check the agent can run itself, with the command named in AGENTS.md. A hook for the moment the agent wants to finish that runs the tests and returns failures to the model as feedback. The result is judged by the exit code and the diff, not the summary. Anthropic and LangChain both describe a premature “done” as one of the most common failures of coding agents. - **Costs doubled after you connected three MCP servers. Why, and what do you do?** Tool definitions sit in the prefix, so you pay for them on every turn, and in a harness that loads them up front, connecting a server mid-session invalidates the cache. What helps: tool search, so only names sit in the prefix, only the servers the task needs, and a fixed tool set for the whole session. Check the effect in usage: the ratio of cache reads to cache writes. - **When is a subagent better than doing the work in the main session?** For side tasks that mostly read: searching the repository, reading logs, reviewing a diff. The subagent spends tokens in its own window and returns a short summary, so the main context stays short. Edits that depend on decisions from the main thread stay in the main session, because the subagent doesn’t see those decisions, and in Claude Code its edits usually don’t go into your checkpoints. In total, subagents cost more tokens, not fewer. ### Sources - [Sebastian Raschka: Components of a Coding Agent (April 2026)](https://magazine.sebastianraschka.com/p/components-of-a-coding-agent) - [Anthropic: Effective harnesses for long-running agents (2025)](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents) - [Claude Code docs: How Claude Code uses prompt caching](https://code.claude.com/docs/en/prompt-caching) - [OpenAI Codex docs: Agent approvals and security](https://learn.chatgpt.com/docs/agent-approvals-security) - [LangChain: Improving Deep Agents with harness engineering (2026)](https://www.langchain.com/blog/improving-deep-agents-with-harness-engineering) --- ## Multiple agents *Agents* *Last edited: 28 September 2026* A lead agent splits the task and delegates parts to subagents, each working in its own clean context and returning a short result. It helps with independent parts, such as researching many sources at once, and hurts with shared state, such as a change to one codebase. It usually costs several times more tokens. **In plain words:** Five researchers will get through ten libraries faster than one, as long as each gets a clear brief and hands back a page of notes. Five programmers fixing the same module without talking to each other will do more damage than one, because each will quietly make different decisions. You pay both teams for every hour. *Interactive widget on the page: Change the number of subagents and the type of task. Watch the time, the tokens, the lead agent’s context and the quality of the result.* ### Lead agent and subagents - A subagent is a separate agent loop (see “The agent loop”) that the lead agent calls like a tool: the argument is a brief, the result is a short report. The subagent does not see the lead agent’s history, only its own system prompt and the brief. It may read tens of thousands of tokens of pages or files, and usually hands back 1–2k. This is the context isolation from “Context engineering and memory”. - The lead agent, also called the orchestrator, plans, hands out the parts, waits and merges the results. In Anthropic’s research system (described in June 2025) it launches subagents in batches and waits for the whole batch. That simplifies coordination, but the slowest subagent sets the pace, and a subagent cannot be corrected while it works. - Other layouts. Pipeline: fixed stages, where one agent’s output is the next one’s input. Handoff: an agent passes the conversation, with its state, to a specialised agent and drops out, for example customer service handing a case to the returns team. Critic loop: one agent writes, another grades the result against criteria and sends back comments until it passes or the round limit runs out. With a subagent the lead agent keeps control; with a handoff it gives control away. ### When it helps - Breadth-first, parallel work: the task splits into independent directions, such as many companies, sources or hypotheses. Anthropic reports that Claude Opus 4 as the lead agent with Claude Sonnet 4 subagents outperformed Claude Opus 4 alone by 90.2% on its internal research eval. Parallel subagents and parallel tool calls cut the time of complex queries by up to 90%. - More tokens per task than one window holds. In the same report, token usage alone explained 80% of the variance in performance on BrowseComp, a benchmark for finding hard-to-locate information. Subagents let you spend those tokens, while only the conclusions reach the lead agent, so its context stays clean. - Specialisation: each subagent has its own prompt and a narrower set of tools and permissions. From a shorter list the model picks the right tool more accurately (see “Tools (function calling)”). A read-only subagent can read untrusted content without access to risky actions, but its report can still carry an injected instruction to the lead agent (see “Prompt injection”). ### When it hurts - Cost and latency. In Anthropic’s data an agent uses about 4 times more tokens than a chat, and a multi-agent system about 15 times more. Each subagent pays for its own system prompt, tools and brief. It starts from zero, so it first gathers context the lead agent already had. - Lost context and duplicated work. A subagent knows only what is in its brief, and every summary loses detail. An early version of Anthropic’s system gave short briefs, such as “research the semiconductor shortage”. One subagent researched the 2021 automotive chip crisis, while two others duplicated each other’s work on 2025 supply chains. - Conflicting decisions in shared state. Every action carries implicit decisions, and parallel agents cannot see each other’s (Cognition, “Don’t Build Multi-Agents”, June 2025). In their Flappy Bird clone example, one subagent built a Super Mario-style background and another a bird in a different style, leaving the lead agent to glue them together. In code this means the same files edited at once, different names and types, merge conflicts. Anthropic itself notes that most coding tasks have fewer truly parallelisable parts than research. - Errors compound, and debugging is hard. A bad split by the lead agent breaks every branch at once, and a small change to its prompt can change subagent behaviour unpredictably. Cemri et al. (2025) analysed more than 1,600 traces from 7 frameworks and described 14 failure modes in three groups: system design, inter-agent misalignment and task verification. On popular benchmarks, multi-agent systems often gained very little. ### How to build it - Start with one agent with tools; OpenAI’s guide gives the same advice. Add more agents once you show, on the same eval set, a gain worth the extra tokens (see “Evals”). For code, a read-only subagent that searches the repository and answers a specific question is usually enough, while a single agent makes the changes. The exception is a large change that splits into independent units, such as a module-by-module migration: each writing subagent works in its own git worktree or sandbox on its own branch and runs the tests, and its change is merged like an ordinary PR (Claude Code’s `/batch` works this way). Each unit still has a single writer, the rule Cognition also arrived at in April 2026. - The lead agent’s brief is a specification: the goal, output format, tools and sources, boundaries (what not to do, what the others are doing) and an effort budget. Anthropic wrote scaling rules into the prompt: a simple question gets 1 agent and 3–10 tool calls, a comparison 2–4 subagents with 10–15 calls each, complex research more than 10 subagents. For a code change, the shared contract (names, types and interfaces) is agreed before the work is split. - Pass large results through files or a store, not through messages. The subagent writes an artefact and returns its path with a short description, so the content does not pass through successive summaries. The lead agent also saves its plan outside its context, because on long tasks the window will run out. - Limits are enforced by code, not by the prompt: the number of subagents, nesting depth, a token and time budget per subagent, the number of critic-loop rounds. Without them, early versions of Anthropic’s system spawned 50 subagents for simple queries. Claude Code (as of September 2026) allows three levels of nesting and at most 20 concurrent subagents by default (see “The coding-agent harness”). - One trace for the whole task, with a separate span for each subagent: brief, calls, tokens, result. Only then can you tell whether the split, a subagent or the merge failed. Grade the end state, for example the facts and sources in the report or passing tests, not the lead agent’s summary (see “The agent loop”). ### Check yourself **Question:** When would you build a system with multiple agents, and when would you stick with one? **Short answer:** Start with a single agent. Add more when the task splits into independent parts, such as researching many sources: a lead agent hands them to subagents, and each works in parallel in its own clean context and returns a short summary. The gain is speed, breadth and a clean lead context. The price is tokens, about 15 times a chat in Anthropic’s data, and coordination. With shared state, such as one change across a codebase, agents make conflicting decisions, so there one agent makes the changes. End-state evals decide. ### Follow-up questions - **A subagent doesn’t know what the others decided. How do you deal with that?** Either pass it the full context and the decisions made so far, which eats up the gain from isolation, or agree the shared things before the split: names, types, output format, the scope of each part. If that can’t be settled up front, the parts aren’t independent and one agent will do better. - **The lead agent spawns 50 subagents for a simple question. What do you do?** That was a real bug in an early version of Anthropic’s system. Scaling rules go into the prompt: how many subagents and calls for which type of question. Code adds hard limits on the number of subagents, the depth and the token budget, and evals check the change. - **How do you pass large results between agents?** Through files or a store: the subagent writes an artefact and returns its path with a short description. Details don’t get lost in successive summaries, and the lead agent’s context stays small. - **How is a handoff different from calling a subagent?** A subagent works like a tool: the lead agent waits for the result and keeps control. A handoff passes the conversation, with its state, to another agent, which carries on talking to the user itself. A handoff suits routing a case to a specialist; a subagent suits splitting a task into parts. - **How do you debug and evaluate a multi-agent system?** One trace ID for the whole task and a separate span for each subagent with its brief, calls and tokens. Grade the end state on a fixed set and compare it with a single agent, because a small change to the lead agent’s prompt can change the behaviour of every subagent. ### Sources - [Anthropic: How we built our multi-agent research system](https://www.anthropic.com/engineering/multi-agent-research-system) - [Cognition: Don’t Build Multi-Agents](https://cognition.com/blog/dont-build-multi-agents) - [OpenAI: A practical guide to building agents (PDF)](https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf) - [Cemri et al.: Why Do Multi-Agent LLM Systems Fail? (arXiv, 2025)](https://arxiv.org/abs/2503.13657) - [Cognition: Multi-Agents: What’s Actually Working (April 2026)](https://cognition.com/blog/multi-agents-working) --- ## Hallucinations *Quality and security* *Last edited: 28 September 2026* A model has no “I don’t know” mode: it always writes a plausible continuation, in the same confident tone whether it knows the answer or is guessing. Hallucinations cannot be switched off; they can be reduced and measured. **In plain words:** A student in an oral exam who never says “I don’t know”. When they know the answer, they speak fluently and confidently. When they don’t, they speak just as fluently and confidently, so you can’t tell the difference from the tone. *Interactive widget on the page: Turn on a source in the context and permission to say “I don’t know”, separately and together. Watch which fabrications disappear, which remain and how many answers you lose along the way.* ### Where they come from - The model predicts the next token and always returns one. There is no separate signal for “my knowledge ends here”, and a guessed sentence is as fluent as a true one. A sampled token cannot be taken back: if the answer began with “In 1978”, the rest of the sentence will justify that year (see “The next token”). - Facts that appeared once or never in the training data are stored in the weights weakly or not at all. The model saw the capital of Australia thousands of times, and the opening year of a small museum once at most. Kalai et al. (OpenAI, 2025) show that for facts with no learnable pattern, such as birthdays, the hallucination rate after pretraining is roughly at least the fraction of facts that appeared exactly once in the data. - Post-training does not remove this, because almost all popular benchmarks grade 0/1: “I don’t know” scores zero, the same as a wrong answer. Under such grading abstaining is never optimal, so a model trained and selected for the score learns to guess, like a student on a test with no negative marking. ### Types - Fabricated fact: a wrong date, number or name stated confidently. Fabricated source: a paper title, court ruling, link or quote that does not exist. In 2023 a federal court in New York fined lawyers $5,000 for a filing that cited rulings invented by ChatGPT (Mata v. Avianca). - Unfaithfulness to the source: the document is in the context, yet the answer contradicts it or adds something it does not contain. In RAG this is the typical generation-stage error, and a dangerous one, because the citation makes the answer look verified. Reasoning error: the facts are right, but the arithmetic is wrong or the conclusion does not follow from the premises. - In an agent: a tool call with an invented argument (customer ID, file path, field name), or a report such as “I sent the email” or “the tests pass” when no tool did any such thing. In code: functions, parameters and whole packages that do not exist. An attacker can register a package under a name models tend to invent and wait for an agent to install it. ### What helps, and its limits - A source in the context with citations (see “RAG”) helps most, but only when retrieval finds the passage containing the answer. Add explicit permission to say “I don’t know” and the instruction “first extract verbatim quotes from the source, then answer based only on them”. A quote can be checked with a plain text search. The price is that the model answers fewer questions. - Tools instead of memory: code or a calculator for arithmetic, search and a database for facts, documentation for APIs (see “Tools (function calling)”). A verification step, meaning a separate call that reads the answer together with the sources and strikes out unsupported claims, catches some of the errors but makes mistakes of its own. - Lower temperature is not a cure: temperature 0 picks the most probable token, so a fabricated date simply becomes reproducible. Reasoning before answering (see “Reasoning models”) helps with logic and arithmetic, because the model can check its steps, but it cannot add knowledge that is not in the weights. Fine-tuning is a poor way to add it: the model learns new facts slowly, and as it learns them its tendency to fabricate grows (Gekhman et al., 2024). ### Confidence threshold and measurement *Interactive widget on the page: Move the confidence threshold below which the model says “I don’t know”. Compare the score under 0/1 grading and under grading with a penalty for errors, and find the point where abstaining pays off.* - Choose the threshold by the cost of a mistake and measure it on your own set (see “Evals”): correct, wrong and abstained answers separately, plus questions whose answer is deliberately missing from the sources. Accuracy alone rewards guessing; error rate alone rewards silence. - Model confidence from `logprobs` is not a measure of truth. A model can be confident in an error, confidence is spread across different phrasings of the same answer, and post-training hurts calibration: in the GPT-4 report the base model was well calibrated on MMLU questions, and noticeably worse after post-training. Without sources, consistency is a better signal: sample the same answer several times and check whether the samples mean the same thing (semantic entropy, Farquhar et al., Nature 2024). Guessed facts change between samples; known ones stay put. - Check faithfulness and citations in code. First, that the cited passage exists verbatim in the document. Native citations in APIs, for example in Claude, guarantee a correct pointer into the document, but not that the passage supports the claim. Then, claim by claim, whether the passage supports it: an NLI model or an LLM judge with a yes/no verdict. ### Check yourself **Question:** Where do hallucinations come from, and how do you reduce and measure them in production? **Short answer:** The model always produces a plausible continuation and has no built-in “I don’t know”, so on rare facts it guesses in the same confident tone. Training and benchmarks graded 0/1 reward a hit and give zero for abstaining, so guessing pays. It can’t be switched off, only reduced: sourced context with citations, permission to say “I don’t know”, tools for facts and arithmetic, and checking citations in code. Lower temperature just repeats the same guess. Measure errors, abstentions and faithfulness to sources separately, and set the answer threshold by the cost of a mistake. ### Follow-up questions - **Can you detect a hallucination from logprobs?** Partly. Low confidence can be a signal, but a model can be confident in an error, confidence is spread across different phrasings of the same answer, and post-training hurts calibration. A better signal comes from sampling several answers and checking whether they mean the same thing. - **Why does a RAG system still make things up?** Retrieval always returns something. If the chunk doesn’t contain the answer, the model often fills the gap from its weights and attaches a citation to the nearest source. On top of that, it can distort the chunk right in front of it. - **The agent reports “tests pass”, but they don’t. What do you do?** Trust the state, not the report. Code runs the tests and sends the result back to the model, and a tool, not the model’s claim, checks the task’s completion condition. In evals you grade the end state, and runs with a false report go into the set as new cases. - **How do you measure hallucinations in your own system?** A set of questions with verified answers plus questions whose answer isn’t in the sources. Count correct, wrong and abstained answers separately, and for RAG also faithfulness: whether every claim is supported by the supplied chunk. ### Sources - [Kalai et al.: Why Language Models Hallucinate (arXiv, 2025)](https://arxiv.org/abs/2509.04664) - [Claude docs: Reduce hallucinations](https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-hallucinations) - [Lilian Weng: Extrinsic Hallucinations in LLMs](https://lilianweng.github.io/posts/2024-07-07-hallucination/) --- ## Evals *Quality and security* *Last edited: 28 September 2026* An LLM’s output is random and sensitive to small prompt changes, so “I checked it on three examples” means nothing. An eval is a fixed set of cases with automated grading, run after every change to the prompt, model and tools, plus quality measurement on production traffic. **In plain words:** Tests in CI for code that answers slightly differently every time. Instead of “pass or fail” you count how many times out of five it passed, and some assertions can’t be written as string comparisons, so a second model checks them, once you have checked that model. *Interactive widget on the page: A revised prompt v2 raises the score from 5 to 7 out of 8. Check in the table what that number hides, then run each case five times and see which change was noise.* ### The set and the grading - Start with error analysis, not metrics: read 50–100 traces, note every problem, group the notes into failure types and count them. The most frequent types show what to fix and what to test. Husain and Shankar put 60–80% of the effort here, not into automated checks. - Build the set, the golden set, from those failures, not from imagination: turn each type into a few cases with an expected output or a grading criterion. Anthropic advises starting with 20–50 tasks. Add every production failure as a new case, plus edge cases and questions that have no answer. - With no traffic yet, seed the set with synthetic cases: list the dimensions of a request (user type, intent, difficulty), write a few combinations by hand, have a model expand them and phrase each as a realistic input, then read the traces they produce. Synthetic cases can’t tell you how common a failure is, come out cleaner than real users and miss the quirks of specialised domains and rare languages, so real traces replace them as traffic arrives. - Not every failure needs an automated evaluator. Fix obvious gaps first, such as an instruction missing from the prompt, and automate only failures that persist: an LLM judge costs time to build and validate, and money on every run. - Red-teaming is a deliberate adversarial set: people, or an attacker model, try to break the system with prompt injection, jailbreaks, requests outside its policy and attempts to extract data. It is kept apart from quality evals, reported as an attack success rate and run on every release (see “Prompt injection”). - Grade with the cheapest method that captures the criterion: exact match, a schema or regex, a code check (run it, inspect the state), a model as judge, a human. One criterion per grader and a yes/no verdict instead of a 1–10 scale, because every grader, a model included, applies a scale differently. Skip ready-made metrics such as “helpfulness”, BERTScore or ROUGE: they measure something other than your failures and give false confidence. - Validate an LLM judge against human labels before trusting its numbers. Raw agreement misleads: if 90% of answers are good, a judge that passes everything agrees 90% of the time and catches no failure. So an expert grades 100–200 examples, you measure separately the share of failures the judge catches (TPR) and of good answers it passes (TNR), and you revise the judge prompt on one part of the labels and check it on another. Known biases: the judge prefers the answer shown first, the longer one and one in its own style (Zheng et al., 2023). Pin the judge’s version and recalibrate whenever it changes. - Evaluate the system piece by piece and as a whole. In RAG, grade retrieval (recall@k, the share of relevant chunks in the top k results) separately from generation (faithfulness to the chunks, see “RAG”). For an agent, grade the end state rather than a fixed sequence of steps, plus hard constraints on the trace: no forbidden tool calls or policy violations, limits on turns and cost. In a long conversation or agent run, find the first failure: later ones usually follow from it. ### Noise and repeated runs - The same case passes one time and fails the next, even at temperature 0: server-side computation produces slightly different numbers depending on batch size, and batch size changes with load. Run each case several times and report the pass rate. - A small sample means a lot of noise. With 50 cases and a score around 80%, the 95% confidence interval is about ±11 percentage points, so a 3-point difference between prompts means nothing. Pairwise comparison on the same cases helps, as does looking at the cases whose result changed rather than only at the mean (Miller, “Adding Error Bars to Evals”, 2024). - Two metrics describe an agent. pass@k: does it complete the task at least once in k attempts, which makes sense when the result can be checked and the best one picked, for example code with tests. pass^k: does it complete it in every one of k attempts, which is the reliability a user expects. At 90% single-attempt success, pass@3 is 99.9% but pass^3 only 73%. In τ-bench (2024), GPT-4o scored below 25% pass^8 on retail customer service. ### Offline, CI and production - Offline you keep two kinds of set. Capability evals start with a low score and show whether a change moves you toward the goal. Regression evals stay near 100% and catch something that used to work and stopped. A set that always scores 100% says nothing about progress, so you add harder cases. - The regression set is a gate in CI. Run it on every change to the prompt, model, tool definitions and retrieval configuration, because each of them changes behaviour. On a pull request, a fast subset with code-based grading; the full set with an LLM judge nightly or before a release. The gate blocks when a critical case starts failing or the score drops below a threshold with a margin for noise, and it also watches cost and latency. - In production, a guardrail blocks or fixes an output before the user sees it, so it must be fast and precise; an evaluator measures quality afterwards. There are no expected answers, so you measure differently: a reference-free judge on a sample of traffic (faithfulness to sources, format, safety), behavioural signals (user corrections, retries, escalations to a human, thumbs-up and thumbs-down ratings) and A/B tests or a canary when you change the model. A fixed set run periodically against a pinned version detects changes on the provider’s side (see “Compute and ‘getting dumber’”). Failed production traces flow back into the offline set. - Public leaderboards tell you which model is worth trying, not whether it will work for you: they measure someone else’s task, have often leaked into training data and saturate quickly. ### Check yourself **Question:** How do you check that an LLM system works and that a change hasn’t broken it? **Short answer:** Build a fixed set of cases from real failures: read traces, name the error types and turn each into a test. Grade with the cheapest method that suffices: exact match, code checks, an LLM judge validated against expert labels, humans. Run every case several times because outputs are random, and compare versions pairwise. The regression set gates every prompt, model and tool change in CI. In production, grade a sample of traffic with a reference-free judge and watch user signals, and feed failures back into the set. ### Follow-up questions - **How do you start without data?** With error analysis on 50–100 real examples: read the outputs, name the error types and turn them into test cases. 20–50 cases are enough to start, as long as they come from real failures. - **When can you trust a model as a judge?** When its grades agree with an expert’s on a held-out sample, measured separately for good and bad answers. A judge that passes everything also scores high accuracy when most answers are good. Give it one criterion and a yes/no verdict, and when comparing two answers, swap their order. - **How do you set up a CI gate that doesn’t block on noise?** Keep stable regression cases that score close to 100% in the gate, and run each several times. Block when a critical case clearly drops or the mean drops by more than the noise margin. Move cases that flicker without any change into quarantine and investigate them instead of loosening the threshold. - **How do you evaluate an agent?** With a test of the end state, not the agent’s summary, running each task several times, with pass^k as the reliability metric. Add cost, number of turns and time per task. Traces show where the agent goes astray. You don’t grade the path rigidly, because several routes lead to the goal, but you do check it for forbidden calls and policy violations. - **You’re switching to a cheaper model. How do you decide?** The same set on both models, pairwise and with repeats, with a separate score for each case type, plus cost and latency. Then a canary or shadow run on part of the traffic and a comparison of production signals. The prompt often needs retuning for the new model, so you compare the best version of each. ### Sources - [Hamel Husain: Your AI Product Needs Evals](https://hamel.dev/blog/posts/evals/) - [Hamel Husain, Shreya Shankar: AI Evals, Everything You Need to Know (synthetic data, judge validation)](https://hamel.dev/blog/posts/evals-faq/) - [Anthropic: Demystifying evals for AI agents (2026)](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) - [Evan Miller: Adding Error Bars to Evals (arXiv, 2024)](https://arxiv.org/abs/2411.00640) - [Zheng et al.: Judging LLM-as-a-Judge (arXiv, 2023)](https://arxiv.org/abs/2306.05685) --- ## Prompt injection *Quality and security* *Last edited: 28 September 2026* To a model, everything is one stream of tokens. There is no separate channel for instructions and another for data, so text the agent merely reads can start steering it. You don’t patch this with a prompt, only with architecture. **In plain words:** An assistant who carries out every instruction they read, including one added in small print to a letter from a stranger. A note saying “don’t follow instructions in letters” won’t help much, because the assistant reads it the same way as the letter. What helps is not giving them the keys to the safe. *Interactive widget on the page: An email agent receives a message from a stranger. Turn off its capabilities one at a time and see whether the hidden instruction still steals data, then make the user the attacker and compare which defences apply.* ### Where it comes from - The system prompt, the user’s message and the email body are the same token stream to the model, separated only by role markers (see “How a model sees a chat”). The model is trained to give priority to instructions higher in the hierarchy, but that is a learned tendency, not a boundary. Against SQL injection you have parameterised queries, which separate code from data. An LLM has no equivalent. - Direct injection: the user attacks, for example by trying to extract the system prompt or get around the rules. Indirect injection: the instruction arrives in content the agent reads on someone’s behalf. The victim is the user, and the attacker needs no access at all; it is enough for the agent to read their text. - Untrusted content is not only email: web pages, PDFs, issues and comments in a repository, tool results, tool descriptions from third-party MCP servers. The instruction can be invisible to a human: white text, an HTML comment, image alt text, Unicode characters that do not show on screen. ### Jailbreak vs prompt injection - A jailbreak is the user getting the model to do what its safety training refuses, through role-play, a harmful request split into innocent parts, base64, or an optimised gibberish suffix that transfers between models (Zou et al., 2023). The attacker is the user and the target is the model’s refusals; in prompt injection the attacker is a third party and the target is the user’s data and actions. OWASP files jailbreaks under direct injection, but for the design what matters is who attacks whom. - Safety training makes a refusal likely for requests that resemble harmful ones from training. It is a learned tendency, not a boundary: it fails on requests unlike its training data (Wei et al., 2023), fine-tuning can undo it (see “Fine-tuning and LoRA”), and it does nothing against injection, because “forward the bank emails to this address” is not a harmful request. From the user, the same sentence would be a legitimate task. - So the defences differ. Against a jailbreak: safety training and classifiers, plus rate limits and account monitoring, because the attacker is a user with unlimited attempts. If the user has fewer permissions than the system, as with a public bot with database access, tools check the user’s own permissions, so a jailbreak gains nothing. Against injection: the architecture below. ### The lethal trifecta - Data theft needs three things at once: access to private data, exposure to untrusted content and a way to send something out. Simon Willison called this the “lethal trifecta” (2025). Remove one and the path to a leak closes. Other harms remain: an agent with write access can delete something, and a summary can lie. - The outbound channel is not only sending email. A URL fetch with data in a parameter is enough, or a markdown image the app downloads on its own, a link someone clicks, a comment in a public repository. EchoLeak (June 2025, Microsoft 365 Copilot, CVSS 9.3): a single email, no click from the victim, an attack classifier bypassed, data exfiltrated through an automatically loaded image fetched via a Microsoft Teams proxy that the CSP allowed. An allowlisted domain that proxies, redirects or hosts user content reopens the channel. - Meta framed it as the Rule of Two (2025): within one session an agent may have at most two of three properties: it processes untrusted input, it has access to sensitive data or systems, it changes state or communicates externally. When all three are needed, it works under human supervision, not on its own. ### Why a prompt won’t fix it - Defensive instructions (“ignore instructions in the data”), delimiters around content and attack classifiers raise the bar, but they work statistically. The attacker keeps trying, and one success is enough. Nasr et al. (2025, with authors from OpenAI, Anthropic and Google DeepMind among others) broke 12 published defences against jailbreaks and prompt injection with attacks adapted to each defence, in most cases with success rates above 90%, although the defences’ authors had reported attack success rates near zero. As Willison puts it, in security a filter that stops 95% of attacks is a failing grade. - The output of a model that has read untrusted content is untrusted too. Don’t render images from arbitrary domains in it, don’t open its links automatically, and don’t pass it unchecked to a privileged tool. ### Defence in depth - Least privilege: per-user access tokens, read-only wherever possible, no secrets in the context. The tool checks authorisation on its own side, not the model. A human approves sensitive actions, and code shows them exactly what is being done and to whom, because an “Approve” button clicked a hundred times a day stops protecting anything. - Isolation: untrusted content is read by a separate model with no tools that returns only structured output, for example a category and an amount. The privileged model never sees the raw text (the Dual LLM pattern). Another pattern, plan-then-execute: the agent fixes its list of calls before it reads untrusted data, so that data cannot change what gets called. CaMeL (Google DeepMind, 2025) separates control flow from data flow and tracks where every value came from: on the AgentDojo benchmark it completed 77% of tasks with provable security, versus 84% with no defence at all. - Output control: an allowlist of domains and recipients, a CSP policy for images, a sandbox without network access for code. Add a log of every tool call and a set of attacks in your evals, run on every release (see “Evals”). Assume some attacks will get through, and design so they can do little. ### Check yourself **Question:** What is prompt injection, and how would you secure an agent that reads emails and web pages? **Short answer:** A model cannot tell instructions from data because everything is one token stream, so content the agent reads can steer it. A leak needs three things at once: private data, untrusted content and an outbound channel. Defensive prompts and classifiers work statistically, and attacks adapted to the defence get through. So design as if the attack succeeds: break that trio within a session, grant least privilege, have code require human approval for sensitive actions, let a tool-less model read untrusted content, and restrict egress to an allowlist. ### Follow-up questions - **Can you fix this with a better prompt or a classifier?** No. They raise the bar but work statistically, and an attacker keeps trying and adapts the attack to the defence. Only architecture gives guarantees: permissions, isolation, output control, approval enforced in code. - **Direct or indirect injection: which is more dangerous?** Direct injection comes from the user, so the risk is whatever the system gives them access to. Indirect injection arrives in content the agent reads on the victim’s behalf: a web page, an email, a PDF, a tool result. It is usually more dangerous, because the attacker needs no access and the harm falls on an unsuspecting user. - **The agent has to read external emails and reply to them. How do you design it?** External emails are read by a model with no tools that returns only schema fields, such as intent and order number. The privileged agent works on those fields, replies only to the thread’s sender and has no access to other mailboxes, and a human approves replies with attachments or to new recipients. - **How can data leak if the agent has no tool for sending?** Through a markdown image with data in its URL that the app fetches on its own, a link someone clicks, a URL-fetching tool or a write to a public place. You block this with a domain allowlist, CSP and no automatic rendering. - **How do you test resistance to injection?** A set of attacks, both direct and hidden in data, goes into the evals and runs on every release. You measure the attack success rate and what an attack could have done. The rate is never zero, so the test also checks whether the architecture limits the damage. ### Sources - [Simon Willison: The lethal trifecta for AI agents (2025)](https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/) - [OWASP GenAI: LLM01 Prompt Injection](https://genai.owasp.org/llmrisk/llm01-prompt-injection/) - [Beurer-Kellner et al.: Design Patterns for Securing LLM Agents against Prompt Injections (2025)](https://arxiv.org/abs/2506.08837) - [Nasr et al.: The Attacker Moves Second (2025)](https://arxiv.org/abs/2510.09023) - [Lilian Weng: Adversarial Attacks on LLMs (2023)](https://lilianweng.github.io/posts/2023-10-25-adv-attack-llm/) --- ## LLMs in production *Production and serving* *Last edited: 28 September 2026* A model call is a remote dependency that is sometimes overloaded, slow and expensive, and sometimes returns something other than what you asked for. A production system is the code around that call: retries and fallbacks, control of latency and cost, traces, validation and safe rollout of changes. **In plain words:** A restaurant with one brilliant but temperamental chef: sometimes swamped with orders, sometimes slow, sometimes sending out the wrong dish. The floor manager can’t fix the chef, but can put the order in again a moment later, keep a stand-in chef ready, check the plate before it goes out and note what went wrong. *Interactive widget on the page: Pick a failure and turn on safeguards. See what the user gets, how long it takes and what it costs.* ### Failures: what to retry and how - Retry only transient errors: 429 (rate limit), overload (529 at Anthropic, 503 at OpenAI), other 5xx, dropped connections and timeouts. A 400, 401 or 403 will fail the same way the second time. So will a 429 from your tier’s monthly spend cap: at Anthropic it has no `retry-after` header, and `error.details.error_code` is `enforced_spend_limit_reached`. - Back off exponentially with random jitter: about 1, 2, 4 s plus a random extra, and when the server sends `Retry-After`, at least that long. Without jitter all clients come back in the same second and knock the overloaded service down again. Cap the number of attempts and the total time. Retry in one layer only: the official SDKs retry on their own (Anthropic’s twice by default), so your own three-attempt loop on top of the SDK turns one click into up to 9 requests. - A retry is safe only when repeating the call breaks nothing. The model call itself changes nothing, but an agent’s tool does. Every action with a side effect gets an idempotency key (for example the task ID plus the step number), and the tool rejects duplicates, so a retried step won’t send a second email. Don’t automatically repeat a response you have already started streaming to the user. - Rate limits are counted in requests and tokens per minute (RPM and TPM; at Anthropic, input and output separately). Anthropic replenishes them continuously with a token bucket, and a 60 RPM limit can behave like 1 request per second, so a sudden burst of traffic gets 429s despite headroom on the per-minute scale. OpenAI counts `max_tokens` towards the limit if it exceeds the request’s estimate. Anthropic counts the tokens actually generated and, for most models, does not count cache reads at all. The limit is shared by the whole organisation, so you need your own queue and per-user limits to stop one customer from blocking everyone else. - The closest fallback is the same model on a platform someone else runs: Claude on Bedrock or Google Cloud, GPT on Azure or Bedrock (as of September 2026). It is a separate account with its own quotas and model IDs, some features arrive there later, and the cache is cold. A different model also saves availability, but a prompt tuned for the primary behaves differently on it and the tool format may differ. Either way the fallback path needs its own evals, because a rarely used path breaks silently. A circuit breaker, after a run of errors, briefly stops calling the failing provider and sends traffic straight to the fallback, letting a trial request through every so often. When nothing works, degrade gracefully: a cached result, a simpler answer or a clear message instead of a spinner that never stops. ### Latency - Two numbers matter: time to first token (TTFT, which is queueing plus prompt processing) and total time (TTFT plus the number of output tokens times the time per token). Streaming shortens only the perceived wait: the user reads from the first token, but the whole answer takes just as long. With streaming, set the timeout on silence between chunks, not on the whole response, and treat an error event mid-stream (it can arrive after 200 OK) as a failure, not as the end of the answer. A long response without streaming risks the idle connection being dropped. The price of streaming is validation: you can check JSON only once it is complete. - Latency grows mainly with output length, because every token is a separate decoding step (see “Why the GPU is idle”). By OpenAI’s rule of thumb, halving the output roughly halves latency, while halving the prompt cuts it by only 1–5%. Reasoning tokens are output too. A concise format, short field names and a lower reasoning effort where it isn’t needed all help. - Run independent calls in parallel, for example question classification and retrieval. Give simple steps such as routing, classification or extraction to a smaller, faster model. Whatever a rule or plain code can handle, do without a model. ### Cost - Put the fixed prefix (system prompt, tools, documents) first, because a cache read costs a fraction of the input price, usually 10% at Anthropic (see “Prompt caching”). Run offline jobs such as nightly classification, data enrichment or eval runs through the batch API: 50% cheaper at OpenAI and Anthropic, with results within 24 h (as of September 2026). - Cascade: a cheap model first, and the expensive one only when a validator rejects the result or confidence is low. It pays off when most traffic is simple and you can cheaply check whether the cheap model coped. Without such a check it is just a quality downgrade. - Hard limits: `max_tokens` on every call (it is a time limit too), a maximum number of agent steps, a daily budget per user or customer, and a cost alert. Store the results of repeatable tasks in an ordinary cache keyed on the input, prompt version and model version. A semantic cache, which matches similar questions, can return the answer to a different question. ### Observability - Every task is a trace, and every model call and tool step is a span within it: input, output, tokens (including cached ones), TTFT and total time, cost, the exact model version from the response, the prompt version and the provider’s request ID. OpenTelemetry has GenAI conventions for this (`gen_ai.*` attributes, still in development). - Metrics: p50 and p95 latency (the mean hides the tail), error rate by type (429, 5xx, timeout, validation), fallback share, cost per completed task rather than per call, and online eval results: an LLM judge on a sample of traffic and user ratings. - Prompts and responses contain personal data and company secrets. Mask PII before storing them, and restrict access to logs and how long they are kept. In the OpenTelemetry conventions, capturing message content is off by default. Turn traces where the system failed into test cases (see “Evals”), so every incident becomes a regression test. ### Guardrails and rolling out changes - Check output in code: schema, types and business rules (the amount is not larger than the balance, the ID exists). On failure, send the model a specific error message and retry once or twice, then fall back or return an error. Strict mode removes syntax errors, not bad values (see “Enforcing output format”). Where the domain requires it, add an input length limit, moderation and PII masking. Risky actions such as payments, deletions or sending anything externally are approved by a human, and code enforces that, not the prompt. - Version prompts like code and pin the exact model version, not an alias that can move. Every change to the prompt, model or parameters goes through offline evals first, then a canary on a few percent of traffic with a metric comparison and a quick rollback. Providers retire old versions, so a migration will come anyway, and your own evals are the only proof that the new version is not worse (see “Compute and ‘getting dumber’”). - Know where the data goes. At OpenAI and Anthropic, content sent through the API is retained by default for up to 30 days (abuse monitoring), and for less only under a zero data retention (ZDR) agreement. At Anthropic, Claude Fable and Mythos 5.x require 30-day retention and are excluded from ZDR, so the model choice is also a data decision. Stateful features such as files, batch results or stored responses keep it longer (as of September 2026). Your logs and tracing tool are a second copy of the same data. - EU AI Act (as of September 2026): since 2 August 2026 a chatbot must tell people they are talking to an AI, and generated content must carry a machine-readable mark (systems already on the market have until 2 December 2026). High-risk uses such as CV screening, credit scoring or grading exams get their obligations from 2 December 2027. ### What the user sees - Show progress and let the user stop it: during reasoning or agent steps, show which step is running (searching, reading a file) next to a stop button. A minute-long spinner looks like a hang, and a wrong path is cheapest to stop early. - Show where the answer comes from: sources linked to the passages they support, and a plain “nothing found” instead of an answer from memory. Signal uncertainty with what you can check (no sources, a failed validator, samples that disagree), not with the confidence the model states about itself (see “Hallucinations”). - Agent actions get a preview before and an undo after: a draft instead of a sent email, soft delete, a change log with a revert. Only what cannot be undone needs the human approval described above. ### Check yourself **Question:** You are shipping an LLM-based feature to production. What do you build around the model call itself? **Short answer:** Treat the model as an unreliable, slow and expensive external dependency. Every call has a timeout, and transient errors (429, overload, 5xx) are retried with exponential backoff and jitter, in one layer only. Tool actions carry idempotency keys. The fallback, ideally the same model on another platform, gets its own evals and a circuit breaker. Streaming, shorter outputs, prompt caching, a cheaper model for simple steps and offline batch jobs cut latency and cost. Every call lands in a trace with tokens, cost and version. Outputs are validated in code, and prompt and model changes ship through evals and a canary. ### Follow-up questions - **The provider has been returning 529 for ten minutes. What does the user see?** After a run of errors the circuit breaker stops calling the provider and traffic goes straight to the fallback, so the user gets the backup model’s answer without waiting through more retries. Without a fallback: a fast, clear message, and tasks that can wait go into a queue. When trial requests succeed, the breaker gradually restores traffic. - **Why not retry every error?** A 400 or 401 will fail the same way the second time, and a 429 from an exhausted spend limit won’t clear on its own. Blind retries multiply traffic at the worst possible moment: under overload, every client making three attempts triples the load. Hence jitter, a limit on total time and retries in one layer only. - **How would you halve the bill without losing quality?** Measure first: cost per completed task, broken down into input, output and cache. Then a stable prefix for prompt caching, batch for everything non-interactive, shorter outputs, a cheaper model where evals show no difference, and a cache for the results of repeatable tasks. Check every change against evals, because a cheaper model can quietly lower quality. - **How do you reconcile streaming with JSON validation?** Full validation is possible only after the last token. In the UI, stream text or parse the structure incrementally and show the fields that are already complete. Where an error is costly, for example before a write or an action, buffer and validate the whole thing. - **How do you move safely to a newer model?** Offline evals on a set built from real cases, a comparison of cost and latency, then a canary on a few percent of traffic with the same metrics and a rollback ready. The prompt usually needs retuning, because a new model reads the same instructions differently. Test a fallback to a different model the same way, because that is a model change too. ### Sources - [OpenAI: Rate limits (retrying with exponential backoff and jitter)](https://developers.openai.com/api/docs/guides/rate-limits) - [Claude docs: Errors (which errors to retry)](https://platform.claude.com/docs/en/api/errors) - [OpenAI: Latency optimization](https://developers.openai.com/api/docs/guides/latency-optimization) - [OpenTelemetry: GenAI semantic conventions](https://github.com/open-telemetry/semantic-conventions-genai) - [EUR-Lex: Digital Omnibus on AI, Regulation (EU) 2026/1744 (AI Act dates)](https://eur-lex.europa.eu/eli/reg/2026/1744/oj/eng) --- ## Why the GPU is idle *Production and serving* *Last edited: 28 September 2026* During text generation the GPU spends most of each step waiting for the weights and the KV cache to arrive from memory, and computes for only a fraction of that time. This one effect explains batching, quantisation, speculative decoding, and why output tokens cost more than input tokens. **In plain words:** A lightning-fast chef has to run to a storeroom at the far end of the building for every ingredient. Cooking for one person, he mostly runs. Cooking for a hundred at once, he runs just as much but cooks a hundred times more. *Interactive widget on the page: Grow the batch and watch when compute catches up with waiting for memory. Then switch to prefill.* ### The per-token arithmetic - A decode step produces one token per conversation but reads every weight from HBM. An 8B model in BF16 is 16 GB, which takes about 4.8 ms to read at 3.35 TB/s. That caps a single conversation at about 210 tokens per second, regardless of the GPU’s compute power. - Arithmetic intensity is the number of operations per byte read from memory. Multiplying a vector by the weights costs 2 operations (a multiply and an add) per parameter, i.e. per 2 bytes in BF16: about 1 operation per byte for each conversation in the batch. An H100 SXM needs about 295 operations per byte on paper (989 TFLOPS / 3.35 TB/s) and about 120 at a realistic 400 TFLOPS. Counting the weights alone, it takes a batch of roughly 100–300 conversations before compute time catches up with read time. - Prefill processes the whole prompt in one pass: weights read once serve thousands of tokens, so a single long prompt keeps the GPU fully busy. Prefill sets the time to first token (TTFT); decode sets the time between subsequent tokens. ### Where batching stops being free - A batch shares the cost of reading the weights, but not the KV cache: each conversation reads its own, so attention in decode stays memory-bound at any batch size. In an 8B model with GQA (like Llama 3 8B) the cache takes about 0.125 MB per token, so a conversation with 32k tokens of context is 4 GB. Sixteen such conversations read 64 GB of cache and only 16 GB of weights every step. “The generation loop and KV cache” covers the mechanism. - Past the crossover point every extra conversation makes the step longer for everyone: total throughput grows more and more slowly while the time between tokens rises. A provider picks a point on that curve. Interactive traffic needs smaller batches and fast tokens; batch traffic can be packed tightly. - In an MoE model one token reads only the experts it was routed to, but the tokens in a batch spread across different experts, so at larger batch sizes you read almost the whole model (see “Mixture of Experts”). ### What it means for the system - Single-stream speed ≈ memory bandwidth / bytes read per token (active weights plus KV cache). That is why quantising weights to 8 or 4 bits speeds up decode almost proportionally (see “Quantisation”), and why speculative decoding gets several tokens out of one read of the weights (see “Speculative decoding”). - Cost per token is mostly a matter of batch size. Continuous batching and PagedAttention exist to fit the largest possible batch into limited memory (see “Continuous batching”). - Output tokens are priced several times higher than input tokens (typically 4–8× at the large providers). Input goes through prefill in parallel, while every output token is a separate decode step that occupies a slot in the batch. - Prefill and decode have different bottlenecks and get in each other’s way on a shared GPU. Large deployments split them into separate pools (disaggregated serving): DeepSeek-V3 describes prefill on 32-GPU units and decode on 320-GPU units, and vLLM, SGLang and NVIDIA Dynamo support the split. On a single server, chunked prefill softens the conflict. - For decode you choose a GPU by memory bandwidth and capacity, not TFLOPS. The H200 has the same compute as the H100 but 4.8 TB/s and 141 GB of memory, so it generates faster with large models and long contexts. The Blackwell B200 (shipping as of September 2026) has 180 GB and 8 TB/s, but its compute grew just as much (about 280 operations per byte in BF16), so the conclusions here still hold. ### Check yourself **Question:** Why is decode memory-bound, and what does that mean for serving? **Short answer:** Every decode step reads all the weights and the KV cache from memory but computes only one token per conversation: in BF16 that is about 1 operation per byte, while an H100 needs about 300 before compute becomes the bottleneck. The speed ceiling for one conversation is memory bandwidth divided by bytes read per token. Weights read once serve the whole batch, so batching raises throughput almost for free until KV cache memory or the latency budget runs out. Quantisation and speculative decoding speed up decode. Prefill is the opposite: it is compute-bound. ### Follow-up questions - **Estimate the maximum speed of a 70B model in FP8 on one H100.** 70 GB of weights at 3.35 TB/s is about 21 ms per step, so at most about 48 tokens per second per conversation, before you add the KV cache. That leaves about 10 GB of the 80 for the cache, so the batch will be small. In practice such a model is split across 2–4 GPUs (tensor parallelism), which adds up both bandwidth and memory. - **You have an SLO of 50 ms between tokens. How do you size the batch?** Step time is roughly (weights + the KV cache of the whole batch) / memory bandwidth, as long as compute fits inside it. Grow the batch until the step reaches 50 ms with headroom for p99, or until you run out of cache memory. Beyond that point throughput only grows at the expense of everyone’s latency. - **Why split prefill and decode onto separate machines?** Prefill is compute-bound and arrives in bursts; decode is memory-bound and runs continuously. On a shared GPU a long prefill stalls other users’ generation, while separate pools can be scaled and sharded independently. The price is transferring the KV cache from the prefill pool to the decode pool. - **Compute utilisation in decode is a few per cent. Is that a problem?** Not by definition. At small batch sizes the right metric is memory bandwidth utilisation (MBU): bytes of weights and cache per step divided by step time and peak bandwidth. Compute utilisation (MFU) is what you optimise in prefill and training. ### Sources - [How To Scale Your Model: All About Rooflines](https://jax-ml.github.io/scaling-book/roofline) - [How To Scale Your Model: All About Transformer Inference](https://jax-ml.github.io/scaling-book/inference) - [Databricks: LLM Inference Performance Engineering (MBU)](https://www.databricks.com/blog/llm-inference-performance-engineering-best-practices) - [NVIDIA: H100 specifications](https://www.nvidia.com/en-us/data-center/h100/) --- ## Continuous batching *Production and serving* *Last edited: 28 September 2026* A server groups conversations into a batch so that one read of the weights serves many users at once. Continuous batching swaps conversations in and out of the batch at every step, and PagedAttention hands out KV cache memory in small blocks, so the same GPU fits several times more of them. **In plain words:** A bus that doesn’t wait for everyone to reach the terminus: whoever gets off frees a seat for the next passenger at the very next stop. And nobody reserves a seat for the whole route in advance; you take one when you board. *Interactive widget on the page: Play both modes and compare how many cells sit empty and at which step the last conversation finishes.* ### How the scheduler works - A static batch fixes its membership once: everyone waits for the longest answer, and new requests wait for the whole group to finish. You don’t know answer lengths in advance, so empty slots are the rule, not the exception. - Continuous batching (iteration-level scheduling in the literature, from the Orca system, OSDI 2022) decides the batch membership at every decode step. A finished conversation frees its slot immediately and a waiting one joins on the next step. It works because each step is a separate forward pass, and the layers other than attention don’t depend on sequence length. - A new conversation starts with the prefill of its prompt, which briefly takes the GPU away from conversations that are mid-generation. Chunked prefill cuts a long prompt into chunks that fit the per-step token budget. vLLM V1 enables it by default: each step it schedules the decoding conversations first and fills the remaining budget with prefill. *Interactive widget on the page: Switch the memory allocation scheme and count how many conversations fit in the same 32 blocks.* ### PagedAttention: KV cache memory - Reserving a contiguous region for the maximum answer length wastes memory on headroom and fragmentation. The vLLM authors measured 60–80% of KV cache memory wasted in the systems of the time. - PagedAttention splits the KV cache into blocks of a dozen or so tokens (usually 16), allocated as generation proceeds and addressed through a block table, like pages of virtual memory in an operating system. Only the tail of the last block stays empty, and waste drops below 4%. More conversations in memory means a bigger batch and a lower cost per token. - Blocks can be shared with copy-on-write. A common prefix, such as a system prompt or several samples of the same answer, is stored in memory once. Server-side prefix caching builds on this (see “Prompt caching”). With several replicas it pays off only if the router sends requests with the same prefix to the replica that already holds it (prefix-aware routing, e.g. in llm-d or NVIDIA Dynamo). - The price of on-demand allocation is that memory can run out mid-generation. The server then preempts a conversation and later recomputes its KV cache from scratch. In vLLM V1 this is the default, rather than swapping the cache out to CPU memory. The user sees it as a sudden stall in the stream. ### What it means in production - Batch size is limited by KV cache memory (number of conversations × context length) and by the latency budget between tokens. Server settings such as `max_num_seqs` and `max_num_batched_tokens` in vLLM set this trade-off directly. - Throughput comes at the cost of latency. Under load, time to first token grows because of the queue and prefills, and time between tokens grows because of larger batches. Measure the two separately, at p50 and p99 (see “LLMs in production”). - Providers’ batch APIs sell a loose SLO: at Anthropic it is 50% cheaper, most batches finish within an hour, and the limit is 24 hours. The provider can use this traffic to fill free batch slots and traffic troughs. For evals and bulk processing it is the simplest saving. ### Check yourself **Question:** How does an LLM server handle many users at once, and what do continuous batching and PagedAttention give you? **Short answer:** A batch shares the cost of reading the weights, so a server wants it as large as possible. Continuous batching decides batch membership at every decode step rather than per request: a finished conversation frees its slot immediately and a new one joins on the next step, so the GPU never waits for the longest answer. PagedAttention allocates the KV cache in blocks on demand instead of reserving the maximum up front: memory waste drops from 60–80% to a few per cent, more conversations fit, and shared prefixes are stored once. The price is preemption when memory runs out. ### Follow-up questions - **What limits the batch size?** KV cache memory for all conversations and the latency budget between tokens. With long contexts memory runs out first; with short ones, the latency budget. - **Users report rare stalls of a few seconds in the token stream. What do you check?** Preemptions caused by running out of KV cache memory (a counter in the server metrics), and long prefills of new requests without chunked prefill. Fixes: a lower limit on concurrent conversations, more memory for the cache, or separate replicas for very long prompts. - **What does PagedAttention give you beyond tighter packing?** Block sharing with copy-on-write. A prefix common to many conversations is stored once, and several samples of the same answer (n > 1, beam search) share the prompt’s blocks. The vLLM authors report up to 55% less memory with such sampling. - **Why does the same prompt at temperature 0 sometimes produce different text?** A kernel’s output depends on the batch size, because the order of floating-point additions changes. Batch size depends on server load, so small differences in the logits sometimes change the chosen token. Thinking Machines (2025) identifies this as the main cause of nondeterminism in inference endpoints. The fix is batch-invariant kernels: vLLM has a Batch Invariance mode (beta as of September 2026) whose output doesn’t depend on batch size or request order, which can cost some speed. It is useful for evals, debugging and RL rollouts. ### Sources - [Kwon et al.: Efficient Memory Management for LLM Serving with PagedAttention (SOSP 2023)](https://arxiv.org/abs/2309.06180) - [vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention](https://vllm.ai/blog/2023-06-20-vllm) - [vLLM: Optimization and Tuning (preemption, chunked prefill)](https://docs.vllm.ai/en/latest/configuration/optimization.html) - [Agrawal et al.: Sarathi-Serve, chunked prefill (OSDI 2024)](https://arxiv.org/abs/2403.02310) --- ## Speculative decoding *Production and serving* *Last edited: 28 September 2026* A cheap mechanism guesses the next few tokens, and the large model checks all of them in one pass, keeping the ones it would have chosen itself. The text is the same as without speculation, but one read of the weights yields several tokens. **In plain words:** An intern writes a draft; the boss reads it and corrects the first wrong word, then the intern carries on writing. Reading is much faster than writing, so together they finish sooner, and the text is exactly what the boss would have written alone. *Interactive widget on the page: Step through three rounds and watch how many tokens one verification accepts and how many large-model passes you save.* ### How verification works - The draft, i.e. the cheap guessing mechanism, proposes γ tokens (usually a few). In one pass the large model computes the distribution at each of those positions, just as in prefill. It keeps the longest matching prefix, inserts its own token at the first mismatch, and adds one more token if everything matches. So every round yields at least one token. - Verification is cheap because decode is memory-bound: a pass over five positions reads the same weights as a pass over one and takes almost the same time (see “Why the GPU is idle”). - When sampling, a proposal x is accepted with probability min(1, p(x)/q(x)), where p is the large model’s distribution and q the small model’s. After a rejection you sample from max(0, p − q), normalised. The result has exactly the large model’s distribution (Leviathan et al.; Chen et al., 2023), and with greedy decoding the text is identical, up to small numerical differences. ### What the gain depends on - With acceptance rate α and γ proposals, a round yields on average (1 − α^(γ+1)) / (1 − α) tokens. For α = 0.8 and γ = 4 that is about 3.4 tokens per large-model pass. The time saving is smaller, because the draft costs something too, and rejected proposals are wasted work. - Acceptance depends on the text. Code, JSON, repetition and low temperature give a high α; creative text at high temperature a low one. Leviathan et al. measured a 2–3× speed-up on T5-XXL. DeepSeek-V3 reports 85–90% acceptance of the second token from its MTP module and 1.8× more tokens per second. - The gain shrinks as the batch grows. The weights are read once for the whole batch anyway, so once a step becomes compute-bound the extra positions cost real compute: with large batches and short contexts, speculation can turn into a loss. The KV cache is different, because every conversation reads its own. With long contexts decode stays memory-bound even at large batch sizes, and a draft that is cheap on cache reads still raises throughput (MagicDec: up to 2.5× for Llama 3.1 8B at batch sizes 32–256; EAGLE-3: 1.38× at batch 64 in SGLang). vLLM recommends speculation mainly for low to medium load, but rates EAGLE and MTP as a medium to high gain under heavy load too. ### Variants and choosing in practice - A separate small model from the same family with the same tokeniser: easy to use, but it takes memory and has to be served alongside the large one. - A draft attached to the large model: extra heads that predict the next positions (Medusa), a light layer working on the large model’s hidden states (EAGLE, now EAGLE-3), or an MTP module the model has from training, as in DeepSeek-V3. These usually guess better than a separate small model. The output stays exact only with standard acceptance and an unchanged large model: Medusa-2, which fine-tunes the large model along with the heads, and Medusa’s “typical acceptance” trade exactness for speed. - No model at all: n-gram matching copies the continuation from the prompt (prompt lookup). It works for code editing, RAG and summaries, where the answer repeats parts of the input. - Choose the method and γ from the α measured on your own traffic, and check the gain at the target load, not on a single request. In providers’ APIs speculation usually runs invisibly. The exception is features like Predicted Outputs in the OpenAI API, where you supply the expected text yourself, for example the file before an edit (as of September 2026 only on GPT-4o and GPT-4.1 models and without function calling; rejected predicted tokens are billed as output). ### Check yourself **Question:** How does speculative decoding work, why doesn’t it change the output, and when does it stop helping? **Short answer:** A cheap drafter, such as a small model, extra heads or n-grams copied from the prompt, proposes several tokens. The large model checks them in one pass, because decode is memory-bound and several positions cost about as much as one. It keeps the matching prefix and inserts its own token at the first mismatch. Accepting with probability min(1, p/q) guarantees exactly the large model’s distribution. The gain depends on the acceptance rate: 2–3× on predictable text at small batch sizes, less as the batch grows, and nothing or a loss once the step is compute-bound (large batch, short contexts). ### Follow-up questions - **When does speculation not help, or even hurt?** When acceptance is low: creative text, high temperature, or a draft trained on different data. Also when the step is already compute-bound, i.e. large batches with short contexts: verifying rejected tokens then takes compute away from other conversations. - **Why is the output distribution exactly the large model’s?** Token x is accepted with probability min(1, p(x)/q(x)), and after a rejection you sample from max(0, p − q), normalised. The two paths together give exactly p(x) for every token. In practice only small numerical differences remain, from the different shape of the computation. - **Acceptance is 0.6 with γ = 5. Do you increase γ?** If anything, decrease it. A round yields on average (1 − 0.6⁶) / 0.4 ≈ 2.4 tokens, and with γ = 3 already ≈ 2.2, so the two extra proposals add almost nothing but cost drafting and verification. A better draft gives a bigger gain, for example one fine-tuned on your own traffic. - **Does the draft have to use the same tokeniser as the large model?** In the classic variant, yes, because you compare tokens and distributions over the same vocabulary. That is why the draft comes from the same model family, or the heads are attached to the large model itself. Newer servers relax this: vLLM can pair a draft and a large model with different vocabularies through a token mapping (use_heterogeneous_vocab, as of September 2026 with greedy drafting only). ### Sources - [Leviathan et al.: Fast Inference from Transformers via Speculative Decoding (ICML 2023)](https://arxiv.org/abs/2211.17192) - [Chen et al.: Accelerating LLM Decoding with Speculative Sampling (2023)](https://arxiv.org/abs/2302.01318) - [Li et al.: EAGLE-3 (2025)](https://arxiv.org/abs/2503.01840) - [Sadhukhan et al.: MagicDec, speculation at large batch sizes and long contexts (2024)](https://arxiv.org/abs/2408.11049) - [vLLM: Speculative Decoding](https://docs.vllm.ai/en/latest/features/speculative_decoding/) --- ## Quantisation *Production and serving* *Last edited: 28 September 2026* Quantisation stores the weights, and sometimes the activations and the KV cache too, in fewer bits by rounding them to a coarser grid of values. The model takes 2–4 times less memory, and generation gets faster because decode reads fewer bytes from memory. **In plain words:** A photo saved as a JPEG instead of at full quality: it is several times smaller and looks the same on a phone screen. Artefacts appear only under heavy compression, and first in the fine detail. For a model, the fine detail is long reasoning, code and less common languages. *Interactive widget on the page: Lower the bit count, add an outlier, then turn on a separate scale per block. Watch the rounding error.* *Interactive widget on the page: Pick a model size and see what hardware the weights alone fit on. Change the bit count in the widget above.* ### How it works - Each weight is divided by a scale, rounded to a number from a small range (−8 to 7 in INT4; a symmetric grid, as in the widget, uses −7 to 7, i.e. 15 levels) and stored like that. At compute time it is multiplied back by the scale. The error is rounding noise, and the network tolerates small noise well. - Outliers are the problem. One large value stretches the grid, and the rest of the weights land on a few points around zero. That is why the scale is computed separately for small groups, typically 16–128 weights. In the activations of models from about 6.7 billion parameters up, systematic outliers appear in a few channels (Dettmers et al., LLM.int8(), 2022). - Post-training quantisation (PTQ) works on a finished model in minutes or hours. GPTQ and AWQ use a small sample of calibration data to round where it hurts the layer’s output least. Quantisation-aware training (QAT) copes better with 4 bits and below, but it requires training, so it is usually done by the model’s author. ### Formats and what they speed up - Weight-only quantisation, e.g. W4A16 (weights in 4 bits, activations in 16), reduces the bytes read in decode, so it speeds up generation at small batch sizes. The multiplication still runs in 16 bits, so prefill and large batches gain almost nothing. - FP8 for weights and activations (W8A8) computes on the tensor cores in FP8: an H100 reaches about 1,979 TFLOPS dense in it versus 989 in BF16. So it also speeds up prefill and large batches. Kurtic et al. (2024) measure FP8 W8A8 as practically lossless, INT8 W8A8 at 1–3% loss, and W4A16 as closer to 8 bits than expected. - The new 4-bit formats are floating-point numbers with a scale per block. MXFP4 uses blocks of 32 weights with a power-of-two scale, about 4.25 bits per weight; OpenAI released the expert weights of gpt-oss in it. NVFP4 uses blocks of 16 with an FP8 scale plus an extra per-tensor scale, about 4.5 bits, and is computed in hardware on Blackwell GPUs. - The KV cache gets quantised too. An FP8 cache takes half the memory per token, so with long contexts it frees more room for the batch than squeezing the weights further (see “The generation loop and KV cache”). ### What you lose and how to decide - Rule of thumb, depending on the model and method: 8 bits is practically lossless, 4 bits with a good method is a small loss, 3 bits and below drop off fast. Losses show up first in long reasoning, code, maths and less common languages, and a benchmark average can hide them. - For the same memory, a larger model in 4 bits usually beats a smaller one in 16. After more than 35,000 experiments, Dettmers and Zettlemoyer (2022) found 4 bits almost always optimal for a fixed total number of model bits. - The choice depends on traffic and hardware. On Hopper GPUs such as the H100, Kurtic et al. recommend W4A16 for single requests and small batches and FP8 W8A8 for heavy traffic with continuous batching. On Blackwell, NVFP4 for weights and activations (W4A4) runs on the FP4 tensor cores at 2–3× the FP8 rate, so 4 bits can win under heavy traffic too, at a somewhat larger loss than FP8: Red Hat measures about 99% of BF16 accuracy for 70B+ models and 95–98% for 7–14B ones. Check quality on your own eval set, because the same model can be quantised differently by different hosts (see “Open-weight models” and “Evals”). ### Check yourself **Question:** What does quantisation give you, what do you lose, and which format would you choose? **Short answer:** The weights are stored in fewer bits, from 16 down to 8 or 4, with a scale computed per small block, because single outliers stretch the grid. The model needs 2–4 times less memory, and decode gets faster because it reads fewer bytes. Weight-only quantisation such as W4A16 helps at small batch sizes. FP8 for weights and activations also speeds up the maths, so on Hopper it wins under heavy traffic; on Blackwell, NVFP4 for weights and activations is faster still. 8 bits is practically lossless, 4 bits costs a little, and below that quality drops fast, first in reasoning and code. Your own evals decide. ### Follow-up questions - **PTQ or QAT?** PTQ quantises a finished model in minutes or hours, using a small sample of calibration data (GPTQ, AWQ). QAT simulates quantisation during training, so the model learns to tolerate it and holds its quality better at 4 bits and below. The cost is training, so QAT is usually done by the model’s author. - **Why are activations harder to quantise than weights?** Weights are known in advance and can be carefully rescaled offline. Activations depend on the input and have large outliers in a few channels. Methods that shift the difficulty onto the weights (SmoothQuant) help, as does FP8, which has a wider range than INT8. - **After quantising to 4 bits the benchmarks look fine, but users complain about code. What do you do?** A benchmark average hides losses in long reasoning chains and code. Compare both versions on your own eval set from that area. If the loss is real, go back to 8 bits or keep the sensitive layers at higher precision. - **When does quantising the KV cache give more than quantising the weights?** With long contexts and large batches, when the cache of all conversations takes more memory than the weights. An FP8 cache holds twice as many tokens, i.e. twice as many conversations or twice the context length. ### Sources - [Maarten Grootendorst: A Visual Guide to Quantization](https://newsletter.maartengrootendorst.com/p/a-visual-guide-to-quantization) - [Kurtic et al.: “Give Me BF16 or Give Me Death”? Accuracy-Performance Trade-Offs in LLM Quantization](https://arxiv.org/abs/2411.02355) - [Dettmers and Zettlemoyer: The case for 4-bit precision (2022)](https://arxiv.org/abs/2212.09720) - [NVIDIA: Introducing NVFP4](https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/) - [Red Hat: Accelerating large language models with NVFP4 quantization (2026)](https://developers.redhat.com/articles/2026/02/04/accelerating-large-language-models-nvfp4-quantization) --- ## Mixture of Experts *Production and serving* *Last edited: 28 September 2026* In a Mixture of Experts (MoE) model, the MLP layer is replaced by many “experts”, and a router picks a few of them for each token. The model has a huge number of parameters but computes with only some of them per token: compute like a small model, memory like a large one. **In plain words:** A clinic. Reception sends you to two of its eight doctors. They all have to be in the building, but you pay for only two appointments. When a crowd arrives every doctor has patients, and reception has to make sure the whole queue doesn’t end up outside one door. *Interactive widget on the page: Click the tokens and watch which experts the router picks. At the bottom: how many experts the whole sentence needs.* ### How routing works - The router is a small linear layer. From the token’s vector it computes a score for each expert, picks the top k (usually 2–8), and sums their outputs with weights derived from those scores. The choice is made separately in every MoE layer and for every token. - Experts replace only the MLP. Attention, embeddings and normalisation are shared, which is why Mixtral 8×7B has about 47 billion parameters rather than 56 billion, of which about 13 billion are active per token. Newer models have many more, smaller experts: DeepSeek-V3 has 256 routed experts and one shared expert in each MoE layer, and picks 8. - An “expert” is not a topic specialist. The Mixtral authors found no clear pattern of assignment by the text’s domain; the choice relates more to syntax, e.g. indentation in code consistently goes to the same experts. - Without load balancing the router collapses onto a few favourite experts, and the rest of the parameters barely learn. Hence an auxiliary balancing loss in training (since the first sparse MoE layers, simplified in Switch Transformer) or, in DeepSeek-V3, mainly a per-expert bias adjusted during training and used only to pick experts, plus a tiny per-sequence balancing loss. ### Consequences for serving - Every expert has to sit in memory. DeepSeek-V3 has 671 billion parameters, 37 billion of them active per token. In FP8 that is about 671 GB of weights alone, more than the 640 GB on eight H100s. - With one conversation, decode reads only the active experts, so it is as fast as a model of that active size. With a larger batch, tokens spread to almost every expert: a step reads almost the whole model, and each expert gets only part of the batch. For compute to catch up with reading, the batch has to be larger than for a dense model by roughly the ratio of total experts to selected ones (see “Why the GPU is idle”). - That is why a large MoE is spread across many GPUs by expert (expert parallelism), with an all-to-all exchange of tokens between GPUs in every layer. DeepSeek-V3 decodes on 320-GPU units, one expert per GPU, and duplicates the most frequently picked experts, because the most loaded expert sets the time of the whole step. ### MoE or a dense model - MoE gives more quality per unit of compute and cheaper training. A dense model of the same total size is usually better, but computes many times more per token. Most leading open-weight models since 2025 are MoE, which is why two numbers are quoted: total and active parameters (e.g. Qwen3-235B-A22B). - On hardware with large but slow memory, like a Mac with unified memory, MoE runs surprisingly fast for a single user, because it reads only the active parameters. gpt-oss-120b (117 billion parameters, 5.1 billion active) fits on a single 80 GB GPU thanks to MXFP4. For the same reason a large MoE runs on one consumer GPU if the expert weights stay in CPU RAM and attention runs on the GPU (llama.cpp `--cpu-moe` or `--n-cpu-moe`). - Self-hosting a large MoE pays off only with enough traffic to fill a batch across many GPUs. With less traffic, an API or a smaller model is cheaper (see “Open-weight models”). ### Check yourself **Question:** What is MoE, and what are its consequences for serving? **Short answer:** In MoE the MLP layer is replaced by many experts, and a router picks a few for every token in every layer. Only the active parameters are computed, for example 37 of 671 billion in DeepSeek-V3, so the quality is closer to a large model at the compute of a small one. The price is memory, because every expert must be loaded, and serving: at larger batch sizes tokens hit almost every expert, so it takes much bigger batches, expert parallelism with all-to-all communication between GPUs, and care to keep the expert load balanced. ### Follow-up questions - **An MoE model has 5 billion active parameters. Will it behave like a 5B model?** It computes like 5B, but the whole model has to sit in memory. With one conversation, decode is as fast as a small model. At a medium batch size each step reads almost all the experts, so it costs as much as reading a large model while doing little computation. Only a very large batch brings the cost per token close to a 5B model. - **Why balance expert load during training?** Without it the router collapses onto a few favourite experts, which get better and are picked more and more often, while the rest of the parameters go to waste. In serving, uneven load means the most loaded expert sets the step time. - **How would you spread a large MoE across GPUs?** Attention usually with tensor or data parallelism, experts with expert parallelism: each GPU holds some of the experts, and tokens are exchanged all-to-all in every layer. What matters most is a fast interconnect between GPUs and duplicating the most frequently picked experts. DeepSeek-V3 does this on 320 GPUs for decode. ### Sources - [Hugging Face: Mixture of Experts Explained](https://huggingface.co/blog/moe) - [Jiang et al.: Mixtral of Experts (2024)](https://arxiv.org/abs/2401.04088) - [DeepSeek-AI: DeepSeek-V3 Technical Report (2024)](https://arxiv.org/abs/2412.19437) - [Maarten Grootendorst: A Visual Guide to Mixture of Experts](https://newsletter.maartengrootendorst.com/p/a-visual-guide-to-mixture-of-experts) --- ## Choosing a model *Choosing a model* *Last edited: 28 September 2026* You choose a model for the task, not for the leaderboard: take the cheapest configuration, meaning model, reasoning effort and provider, that passes your eval set within your latency and limits. Public benchmarks only tell you which candidates are worth testing. **In plain words:** Hiring for a specific role. You don’t hire whoever scored highest on a general aptitude test; you give every candidate the same work sample from your actual job and take the one who does it well and on time for the lowest rate. Rankings and diplomas only tell you whom to invite for an interview. The work sample stays in the drawer, so when someone leaves you can vet a replacement in an hour. *Interactive widget on the page: Set a quality floor and a latency cap. See which model is the cheapest of those that pass, then switch to the public leaderboard.* ### From the task to the model - Start from a set of cases taken from real traffic and the bar it has to clear (see “Evals”). Run it on a few candidates, each case several times, and take the cheapest one that passes. - Models come in tiers. Frontier models take hard reasoning, long agentic tasks and code. Mid-tier models cover most production work. Small, fast models handle classification, extraction, routing and simple steps at scale. - Reasoning effort is the second knob of the same choice. In their model-selection guides Anthropic writes that tuning effort is often a better lever than switching models, and OpenAI advises keeping the lightest setting that meets your quality bar (as of September 2026). So you compare model and effort pairs, not bare models. - One model for the whole system is rarely optimal. In a cascade a cheap model answers first, and only what a validator, tests or low confidence reject goes to a stronger model or a higher effort (see “LLMs in production”). Inside an agent you pick a model per step: a small one summarises tool results, a frontier one plans and makes the hard calls. ### Cost and latency - Count cost per completed task, not price per token: all tokens from all attempts, failed ones included, divided by the number of tasks passed. A stronger model often finishes in fewer turns with fewer fixes. In Anthropic’s measurements on a SWE-bench Pro subset, Fable 5.1 at low effort solved 88.6% of tasks for 0.54 USD per solved task, while Sonnet 5, five times cheaper per token, solved 77.4% for 0.84 USD. On long research it went the other way: Fable 5.1 cost about four times more per task. These are vendor numbers, so check them on your own task. - The price list hides three things. Reasoning tokens are billed as output, even when you never see them (see “Reasoning models”). Models differ in verbosity, so on the same task one can generate several times as many tokens as another. The cache discount differs too: at Anthropic a cache read usually costs 10% of the input price, 5% on Opus 5.5 and 2.5% on Fable 5.1 (as of September 2026, see “Prompt caching”). For an agent that reads most of its input from cache, that can flip the comparison. - Latency is two numbers: time to first token (TTFT) and tokens per second after it. Reasoning effort moves both. The first word of the answer arrives only after all the thinking, and useful tokens per second drop because part of what is generated is thinking the user never sees. That is why Artificial Analysis measures time to first answer token separately. A slower model is helped by streaming, lower effort, or moving the step to where nobody is waiting. ### Limits and data - The context window on the model card is an upper bound, not a promise: quality drops long before it is full (see “The context window and agents”), so test at your real lengths. The output limit is separate and much smaller, and thinking counts towards it. - Rate limits depend on your account tier, which grows with your spending history. At Anthropic a new organisation may start on a lower tier, and each tier has a monthly spend cap after which the API returns 429 until the start of the next month (as of September 2026). Check the limits for the exact model before launch, not after. - Strict structured output, parallel tool calls, images, logprobs, batch and prompt caching are not available on every model or at every provider (see “Enforcing output format”). One missing feature can rule out the model that wins on your eval. - Where inference runs and where data is stored are two separate settings. Processing in a chosen region covers only some models and features, may need the provider’s approval, and costs more: about 10% on newer models at OpenAI and Anthropic (as of September 2026). Anthropic’s own API can pin only the US; an EU region comes through Bedrock or Google Cloud. Zero data retention (ZDR) needs an agreement and does not cover stateful features (see “LLMs in production”). ### Reading benchmarks - Tasks from public test sets leak into training data, and once the top models approach the ceiling, differences vanish into noise and into errors in the set itself. In February 2026 OpenAI stopped reporting SWE-bench Verified: in the audited part of the set, 59.4% of tasks had flawed tests that rejected correct solutions, and every frontier model checked had seen some of the tasks and solutions in training. - SWE-bench Verified is 500 bug fixes from 12 Python repositories. It measures small, well-described patches in popular projects, not work in your repository, with your conventions and changes across many files. - LMArena measures which answer a voter prefers, and voters prefer longer, richly formatted ones. Once style was controlled for (style control, 2024), length turned out to be the strongest factor and the ranking reshuffled: GPT-4o-mini fell from 6th to 11th place, while Claude 3 Opus rose from 16th to 10th. It measures chat preference, not correctness on your task. - A vendor-reported score and an independent one are different measurements. The vendor picks the harness, reasoning effort, number of attempts and task subset. Independent evaluators such as Artificial Analysis and Epoch AI run every model in one harness, and Epoch notes that SWE-bench results depend heavily on the scaffold. Compare only numbers from one source and one harness (see “The coding-agent harness”). - The same model at different providers is not the same product: different quantisation, chat template, tool-call parser and serving settings. In Moonshot’s November 2025 measurement, Kimi K2 Thinking produced schema-valid tool calls 100% of the time through the official API and 83–87% of the time at several third-party hosts. Test the specific endpoint, not the model name. ### A choice that survives change - Models are retired fast. Anthropic gives at least 60 days’ notice: Claude Sonnet 4, released in May 2025, got a retirement date in April 2026 and stopped working on 15 June 2026. Pin a specific version instead of an alias, and keep the eval set in your repository, so that switching models is one run plus a canary rather than a project (see “Compute and ‘getting dumber’”). - Route model calls through one thin layer, your own or a gateway, that knows prices, limits and fallbacks. Then a change of model or provider is configuration. You will still have to retune the prompt, because a new model reads the same instructions differently. - Open or closed is a separate decision about control over data, versions and cost at high volume (see “Open-weight models”). And before you pick a model, check whether the task needs an LLM at all (the next page, “When not to use an LLM: classifiers and System One”). ### Check yourself **Question:** How do you choose a model for a new feature, and why doesn’t a public leaderboard settle it? **Short answer:** Start from the task and your own eval set with a quality bar. Run candidates from several tiers, frontier, mid and small, on that set at different reasoning efforts, and take the cheapest configuration that clears the bar within the required latency. Count cost per completed task, including reasoning tokens, the model’s verbosity and the cache discount, rather than reading it off the per-token price list. Then check the limits: effective context length, output limit, rate limits, structured output, tool calling and data region. Public leaderboards can be contaminated, saturated, and sensitive to style and harness, so they only narrow the candidate list. Pin the version and keep the eval set in the repository, so that switching models is a single run. ### Follow-up questions - **The cheapest model scores 81% against an 80% bar. Is that enough?** Not without checking the noise. With 50 cases the 95% confidence interval is about ±11 percentage points, so 81% and 80% are indistinguishable. You need more cases or repetitions, a paired comparison with a more expensive candidate, and a separate score for critical cases. If the cheap model fails on one type of case, that type can be routed to a stronger model. - **Model A costs five times more per token than model B. When does A come out cheaper?** When it finishes the task in fewer steps: fewer agent turns, fewer fixes, shorter thinking at a lower effort, and more tasks passed. Cost per completed task is the cost of all attempts divided by the successful ones, so a model that passes 60% of tasks pays for 40% of its attempts for nothing. A measurement on your own set with full token accounting, cache and reasoning included, settles it. - **How do you build routing between a cheap and an expensive model?** The simplest way is a cascade with a verifier: the cheap model answers, and a schema validator, tests or a judge decide whether to accept the result or rerun the task on a stronger model or at a higher reasoning effort. A router that predicts difficulty up front makes its own mistakes, so you measure its accuracy and the cost of its errors. In Anthropic’s measurements, running everything at low effort and rerunning failures at high effort held the score at about half the cost, but that only works where there is a reliable failure signal. - **A new model scores higher on a public benchmark. Should you switch?** The public score only qualifies it as a candidate. What counts is your own eval set, the same one the current model passed, run pairwise with repetitions, plus cost per task, p95 latency and support for the features you use. The prompt usually needs retuning, so compare the best version for each model. Then a canary on part of the traffic, watching production signals. - **Why put an abstraction layer over the API if you use only one provider?** Because providers retire versions, usually after a year or two, have outages and rate limits, and a better model may appear elsewhere. One layer keeps model names, prices, limits, fallback and token logging in one place, so a switch is configuration plus an eval run. It cannot hide feature differences such as tool formats, thinking or caching, so keep it thin. ### Sources - [Anthropic: Optimizing for cost and intelligence (cost per completed task, effort, cascades)](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence) - [LMSYS: Does style matter? Style control in Chatbot Arena (2024)](https://lmsys.org/blog/2024-08-28-style-control/) - [Epoch AI: SWE-bench Verified benchmark review (with OpenAI’s 2026 audit)](https://epoch.ai/benchmarks/swe-bench-verified/review) - [Artificial Analysis: how TTFT and output speed are measured](https://artificialanalysis.ai/methodology/performance-benchmarking) - [Moonshot AI: K2 Vendor Verifier (one model, many hosts)](https://github.com/MoonshotAI/K2-Vendor-Verifier) --- ## Open-weight models *Choosing a model* *Last edited: 28 September 2026* Open-weight means you can download the weights, i.e. the model’s trained matrices, and run it on your own hardware. It is not the same as open source: you usually don’t get the data or the training code, and the licence may restrict use. **In plain words:** A closed model is a restaurant: you eat what they serve and can’t look into the kitchen. Open-weight is a takeaway: you can reheat and season it at home, but you don’t know the recipe. Open source is the dish with the recipe. *Interactive widget on the page: Switch the model type and see what comes in the package and what follows from it.* ### What you get and what you don’t - Weights, inference code and the tokeniser are enough to run, quantise and fine-tune the model, for example with LoRA (see “Fine-tuning and LoRA”). They are not enough to reproduce it or check what it was trained on. - The OSI’s Open Source AI Definition (version 1.0, 2024) requires the weights, the complete training and inference code, and a description of the data detailed enough for a skilled person to build a similar system, all under OSI-approved terms: use for any purpose without asking permission. The data itself need not be published, but the public and third-party training data must be listed with where to get it. Few models meet it, e.g. Ai2’s OLMo; licences with use restrictions, such as Llama’s, fail it on terms alone. - Licences vary a lot. Apache 2.0 and MIT (e.g. gpt-oss, Qwen3, DeepSeek-R1) allow almost anything. The Llama licences (3.1 and 4) require a separate licence from Meta if your products had more than 700 million monthly active users on the model’s release date, plus a “Built with Llama” notice. For the multimodal models (Llama 3.2 Vision, Llama 4 Scout and Maverick) the acceptable use policy does not grant the licence at all to individuals and companies based in the EU; end users of a product built on them are exempt. ### Why companies want it - Data never leaves your infrastructure: your own cloud, your own data centre, even an air-gapped network. In regulated industries this is sometimes the only acceptable route. - The version is frozen: nobody swaps the weights, settings or serving stack without you knowing (see “Compute and “getting dumber””). This holds only if you host it yourself. - Full control over inference: fine-tuning, your own quantisation, access to logits, any kind of output-format enforcement, no rate limits imposed by a provider. *Interactive widget on the page: Set the API price, the node cost, the monthly volume and what one node can serve. See at what load your own GPUs start to pay off.* ### Cost and pitfalls - A node costs the same at 5% and at 90% load, and traffic has peaks and troughs. Real average load is far below 100%, so you calculate the cost at that load, not at full load. On top of that come people: serving (vLLM, SGLang), monitoring, updates, security. Load weights as safetensors from a trusted publisher at a pinned revision: a pickle checkpoint can run code when loaded, and `trust_remote_code` runs the repository’s Python on your servers. - A third route is the same open model through a third-party host’s API: no GPUs of your own, and usually cheaper than the closed frontier models. Data leaves your infrastructure again, and the host may change the quantisation or the serving stack. - The same model behaves differently at different hosts: different quantisation, a different chat template, different handling of tool calling and long context. Moonshot measured this for Kimi K2 (November 2025): tool-call schema accuracy ranged from 100% on the official API to 72% at one host, because of outdated serving versions, malformed tool-call IDs and no constrained decoding. Test the specific endpoint with your own eval set, not the model name (see “Evals” and “LLMs in production”). - On the hardest tasks the closed frontier models usually lead the open ones. The gap can be small and shifts with every release, so a measurement on your task decides, not a leaderboard. ### Check yourself **Question:** Open-weight vs open source: what’s the difference, and when should you self-host? **Short answer:** Open-weight means you can download the weights and run the model yourself. Open source, per the OSI definition, also requires the training code, a detailed description of the data and terms without use restrictions, which is rare; weights licences often restrict use, for example by scale or region. Self-hosting gives control: data stays in-house, the version never changes, and you can fine-tune and quantise. You pay for GPUs and people, and a node costs the same idle or full, so it pays off at high, steady volume or under hard regulatory constraints. The middle path is an open model from a third-party host. ### Follow-up questions - **What do you check in a licence?** Commercial use, scale thresholds (e.g. 700 million monthly active users in Llama), geographic restrictions (the multimodal Llama models are not licensed to individuals or companies based in the EU), the list of prohibited uses, terms for derivative models and for training other models on the outputs, and attribution requirements. - **How do you compare a host of an open model with the original?** With the same eval set against the specific endpoint, focusing on tool calling, format enforcement and long context, because that is where differences in quantisation and chat template show up first. Also record the quantisation and version the host declares. - **When would you advise against self-hosting despite a lower price per token?** When traffic is irregular and the node sits idle most of the time, when the team has nobody to run the serving stack, or when the task needs the quality of a frontier closed model. Calculate cost at the real average load, not at 100%. - **A client requires that data never leaves the EU. What are your options?** A closed model through a cloud with an EU region and a data processing agreement, an open model at a host in the EU, or self-hosting. The choice depends on whether a contract and a region are enough or full control over the infrastructure is required. ### Sources - [Open Source Initiative: The Open Source AI Definition 1.0](https://opensource.org/ai/open-source-ai-definition) - [Meta: Llama 4 licence](https://github.com/meta-llama/llama-models/blob/main/models/llama4/LICENSE) - [Meta: Llama 4 acceptable use policy](https://github.com/meta-llama/llama-models/blob/main/models/llama4/USE_POLICY.md) - [Moonshot AI: K2 Vendor Verifier (differences between hosts)](https://github.com/MoonshotAI/K2-Vendor-Verifier) - [Ai2: OLMo, a fully open model](https://allenai.org/olmo) --- ## When not to use an LLM: classifiers and System One *Choosing a model* *Last edited: 28 September 2026* When a system needs a decision rather than text (a label, a score, yes or no), an LLM still writes it token by token, and its probabilities are often missing or poorly calibrated. A non-generative model is often faster, cheaper and better calibrated; an LLM wins when the answer space is open or the task needs reasoning. TypeSafe’s System One serves here as the worked example of a new model type. **In plain words:** The difference between writing half a page of justification and ticking boxes on a form. A form needs no writer, but it only has the boxes someone thought of in advance. *Interactive widget on the page: Run both lanes: the LLM writes a three-field JSON token by token, Jev (TypeSafe’s System One model) gets the same three questions in a single call.* ### The options for a decision - A fine-tuned encoder (a BERT-type model) or embeddings with logistic regression: one forward pass, milliseconds, a fraction of a cent, running next to the service, with probabilities you can calibrate on a validation set. The price is labelled data for every task and retraining for every new class. - Zero-shot, with no training data: an NLI model scores each label, written in natural language, as an entailment of the text (Yin et al., 2019), and a cross-encoder scores text–label pairs. A cheap baseline to measure everything else against. - An LLM as a classifier: a prompt with the list of labels and a single-token answer read from logprobs. That is one decode step after prefill, so it is fast too, but each question is a separate request or a longer output. Logprobs are often unavailable (never in the Claude API; in OpenAI’s current models only with reasoning effort set to none, as of September 2026), and post-training degrades calibration: in the GPT-4 report the base model was well calibrated and the post-trained one noticeably worse. Measure calibration and correct it, e.g. with temperature scaling. - An LLM wins when you need text, code, an explanation, multi-step reasoning, arithmetic or tool calls, when the answers cannot be closed into a list up front, and at low volume, where a prompt is cheaper than a labelled dataset. Strict structured output gives an LLM typed answers too (see “Enforcing output format”), so the real differences are cost, latency and the quality of the probabilities. Which LLM to pick is covered in “Choosing a model”. ### Worked example: System One (vendor claims) - According to TypeSafe’s documentation (September 2026, version jev-1.13), its model Jev takes a text “state” (a string, a JSON object or an array of texts) and typed questions: Choice (pick from a list), Score (a level on a described scale) and Noul (the probability of “yes”). Choice and Score return a distribution and a `confidence` field. It generates no text. The vendor positions it as needing no task-specific training, unlike a fine-tuned encoder. - All questions in a request are evaluated in parallel and independently against the same state. The vendor says adding questions barely increases response time but publishes no latency figures. Pricing: $0.042 per million input tokens, output free, up to 64k tokens per request, of which the state plus the longest question may take at most 32k. English is the primary language. - The probabilities were trained for calibration with a method the vendor calls RLCD (reinforcement learning for calibrated decisions), and the vendor cautions that calibration holds across groups of decisions, not for a single answer. The documentation describes neither the architecture nor any independent comparisons, so this whole section is claims to test on your own data. - Limitations the vendor lists for jev-1.13: it does not count reliably, compares dates and numbers poorly, reads questions literally, gets lost with indirect questions and in a large state full of irrelevant data, and the content of the state can steer it the way prompt injection steers an LLM. Answers to a question and to its negation need not sum to 1. *Interactive widget on the page: Move the threshold: decisions above it run automatically, those below go to an LLM or a human. Watch how many mistakes the automation lets through.* ### Fitting it into a system - The pattern: a fast decision model settles the easy cases, and on low confidence the case goes to a reasoning LLM or a human (see “LLMs in production”). Automation pays off when (1 − p) × cost of a mistake < cost of escalation. With a mistake costing $100 and an escalation $5, the threshold is p > 0.95. - Check calibration yourself, whatever the model: group decisions by their stated probability and compare it with the actual accuracy in each group (a reliability diagram, ECE). Do it separately for each question type and language, because calibration in English does not guarantee it in any other language. - Arithmetic, dates and hard rules stay in code. Once thresholds are tuned, pin the model version (for Jev `jev-1.13.0`, not the `jev-latest` alias): an alias moves with each release, and the probability distribution moves with it. ### Check yourself **Question:** When does a classifier or a decision model such as System One beat an LLM, and when doesn’t it? **Short answer:** A non-generative model wins when the answer is a label, a score or yes/no from a closed list and volume, latency or calibrated probabilities matter. A fine-tuned encoder answers in milliseconds but needs labelled data for each task; zero-shot NLI encoders need none. An LLM needs no data but writes the decision token by token, and its logprobs are often unavailable or poorly calibrated after post-training. Decision models such as TypeSafe’s Jev claim typed answers with calibrated probabilities in one call, which has to be measured on your own data. An LLM wins for text, reasoning, tools and open answer spaces. The pattern: the fast model settles confident cases and escalates the rest to an LLM or a human. ### Follow-up questions - **When is an LLM still the better choice?** When you need text, code, an explanation of the decision, multi-step reasoning, arithmetic or tool calls. Also when the space of answers cannot be closed into a list of options up front, or the volume is too low to justify collecting labelled data. - **How do you check calibration?** On your own labelled data: split the decisions into bins by stated probability and in each bin compare it with the actual accuracy. Report the weighted average gap as ECE. Do it separately for each question type and language. - **How does a decision model differ from an LLM classifier with logprobs?** A classifier that emits a single label token is one decode step after prefill, so it is fast too. But each question is a separate request or a longer output, logprobs are often unavailable (never in the Claude API; in OpenAI’s current models only with reasoning off), and post-training does not optimise them for calibration and usually spoils what pretraining gave. Measuring accuracy, calibration, latency and cost on the same dataset settles it. - **How would you set the escalation threshold?** From costs: automate when (1 − p) × cost of a mistake is lower than the cost of escalation. Then check on data that the actual error rate at that threshold matches the expected one, and repeat that after every model version change. ### Sources - [Guo et al.: On Calibration of Modern Neural Networks (2017)](https://arxiv.org/abs/1706.04599) - [Yin, Hay and Roth: Benchmarking Zero-shot Text Classification (2019)](https://arxiv.org/abs/1909.00161) - [TypeSafe: System One (documentation)](https://docs.typesafe.ai/concepts/system-one) - [TypeSafe: Models (pricing, limits, aliases)](https://docs.typesafe.ai/models) - [TypeSafe: Jev 1.13 jaggedness (known limitations)](https://docs.typesafe.ai/model-jaggedness/jev-1.13) --- ## Compute and “getting dumber” *Choosing a model* *Last edited: 28 September 2026* The weights of a pinned model version do not change, but the system that serves them changes all the time: hardware, compilers, the sampler, routing, the app’s system prompt and its defaults. Behind an alias or a model name in a chat app, even the weights can be swapped while the name stays the same. On top of that, the provider has a finite pool of chips that both the training of the next model and today’s users draw from. **In plain words:** One kitchen with two orders at once: a big banquet next month, which is training, and today’s diners, who are the users. The number of burners stays the same. In a pinned version the recipe does not change, but a new oven or the wrong spice will change the taste. *Interactive widget on the page: Move training and demand. See when inference runs out of capacity and what users see when it does.* ### What changes while the weights stay put - The same model runs on different hardware and software stacks. Anthropic says it serves Claude on AWS Trainium, NVIDIA GPUs and Google TPUs. A different precision, compiler or sampler implementation produces slightly different numbers at the output, and a bug in any of those places breaks only part of the traffic. - A capacity shortage shows up first in latency and availability: queues, longer time to first token, slower tokens as batches grow, overload errors, stricter rate limits. It does not lower quality, but load sets the batch size, and with kernels that are not batch-invariant the same request can get different tokens even at temperature 0. - Apps such as a chat or a coding tool change their system prompt, default reasoning effort and the way they manage context. That changes results without changing the model and without any change to the API. ### Documented cases - **April 2025** (OpenAI, “Expanding on what we missed with sycophancy”). Here the weights did change, under the same name. On 25 April an update to GPT-4o in ChatGPT with new post-training, including an extra reward signal from thumbs-up and thumbs-down ratings, made the model noticeably sycophantic; the rollback began on 28 April. OpenAI counts five major post-training updates to GPT-4o in ChatGPT since its launch. - **August and September 2025** (Anthropic postmortem of 17 September 2025). Three overlapping infrastructure bugs, weights unchanged. From 5 August some Sonnet 4 requests went to servers configured for the upcoming 1M-token context, 16% of them at the worst hour on 31 August. From 25 August a misconfiguration of TPU servers caused out-of-place tokens to be inserted, e.g. Thai or Chinese characters in English answers. A change in token selection also exposed a bug in the XLA compiler for TPUs, through which approximate top-k sometimes dropped the most probable token. Fixes landed between 2 and 18 September. - **March and April 2026** (postmortem of 23 April 2026). Three changes in Claude Code, the Claude Agent SDK and Claude Cowork; the API was not affected. On 4 March the default reasoning effort was lowered from high to medium to cut latency; this was reverted on 7 April. On 26 March a bug in a mechanism meant to clear old thinking blocks once after an hour of inactivity made it clear them on every following turn, so the model seemed forgetful and repeated itself; fixed on 10 April. On 16 April a length limit on answers was added to the system prompt, which lowered an eval score by 3%; reverted on 20 April. - According to both Anthropic postmortems, the causes were infrastructure bugs and product decisions made to reduce latency and response length, not a lack of capacity. In 2025 Anthropic stated plainly that it never reduces model quality due to demand, time of day or server load, and in 2026 that it never intentionally degrades its models. ### Hypothesis and illusion - The hypothesis: when chips run short before a launch, the provider quietly serves a more heavily quantised or smaller variant. That is technically possible at any provider, but there is no public evidence for it. Treat it as a hypothesis to test by measurement. - The illusion: expectations rise, tasks get harder, sessions get longer, and a long context lowers quality on its own (see “The context window and agents”). With sampling, individual bad answers always happen, so anecdotes cannot tell a regression from noise. ### How to protect yourself - Pin an exact model snapshot ID, not an alias that can move; for Claude from 4.6 on, the dateless ID is itself the snapshot. Set reasoning effort and the token limit explicitly, and temperature only where the model still accepts it: newer Claude models (from Opus 4.7 on) and OpenAI models with reasoning enabled reject non-default values (see “The next token”). Version the prompt and the tool definitions (see “LLMs in production”). - Run your own eval set on a schedule against the production endpoint, with several repetitions per case, because differences of a few percent get lost in the noise (see “Evals”). Production traces let you compare behaviour before and after. - A pinned version also has a retirement date, so a migration will come anyway; run it against the same eval set. Full reproducibility needs an open model on a frozen stack of your own with deterministic, batch-invariant kernels (see “Open-weight models”). ### Check yourself **Question:** Users say the model has “got dumber”. How would you explain that, and what would you do? **Short answer:** First check whether the model changed at all: in an app or behind an alias the same name can point to new weights, as GPT-4o in ChatGPT did in April 2025. The weights of a pinned version do not change, but everything around them does: hardware and compilers, the sampler, routing, the app’s system prompt, default reasoning effort, rate limits. Both documented regressions at Anthropic, in 2025 and 2026, had such causes, not different weights. Silent quality cuts due to a compute shortage are technically possible but unproven. So instead of guessing, pin the version, set parameters explicitly, run your own evals on a schedule and compare traces. ### Follow-up questions - **How do you tell a regression at the provider from one of your own?** A pinned version, a fixed eval set with several repetitions, and a comparison of traces. If the prompt, tools and parameters are the same and the scores have dropped beyond the noise, the problem is on the provider’s side. Then you collect examples and report them. - **What do you pin?** The exact model snapshot ID, reasoning effort, the token limit, temperature where the model still accepts it, and the versions of the prompt and tool definitions. A pinned version also has a retirement date, so you run the migration against the same eval set. - **Why didn’t the provider catch the 2025 regression itself?** According to the postmortem, Anthropic’s evals did not capture the drop, user reports were noisy, and privacy rules limit engineers’ access to conversations. The bugs hit part of the traffic on some platforms, so the averages looked fine. The lesson for you: you need evals on your own traffic. - **How do you detect a regression within an hour rather than a week?** A canary: a small set of cases run every hour against the production endpoint, plus signals from traffic: the rate of invalid JSON, response length, number of tool calls, retries and user corrections. Alert on any deviation from the baseline. ### Sources - [Anthropic: A postmortem of three recent issues (2025)](https://www.anthropic.com/engineering/a-postmortem-of-three-recent-issues) - [Anthropic: An update on recent Claude Code quality reports (2026)](https://www.anthropic.com/engineering/april-23-postmortem) - [OpenAI: Expanding on what we missed with sycophancy (2025)](https://openai.com/index/expanding-on-sycophancy/) - [Anthropic: Model IDs and versioning](https://platform.claude.com/docs/en/about-claude/models/model-ids-and-versions) - [Horace He: Defeating Nondeterminism in LLM Inference (2025)](https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/) --- ## Design a system *Practice and review* *Last edited: 28 September 2026* Knowing the mechanisms is half the work. The other half is combining them into a system that fits a budget and survives production: from requirements through architecture and trade-offs to evaluation and failures, with numbers rather than the name of a favourite model. **In plain words:** A conversation with an architect about your house. A good one first asks how many of you there are, what your budget is and what plot you have, then sketches the simplest house that meets those conditions, and tells you what happens when the family grows. Whoever starts by choosing the roof tiles, that is, the model, is not designing but picking from a catalogue. *Interactive widget on the page: Pick a scenario and go through it step by step. Before you read a step, decide what you would do, then compare. Each step ends with an open question to settle next.* ### A company-wide LLM gateway - **Requirements.** “Shared LLM access for 40 teams.” Ask: how many providers, whether a self-hosted model is needed (see “Open-weight models”), whether personal data may leave the company, what budgets per team, what availability. Streaming must pass through with no noticeable added latency. *Open question: who pays for the tokens, and who decides which models are allowed?* - **Architecture.** One endpoint in a popular API format, and behind it: a key per team, limits and budgets, PII masking, routing with fallback, a trace of every call. The gateway is stateless and scales horizontally, with budget counters in a shared, fast store. *Open question: how do you version the contract when providers add their own parameters, such as reasoning effort?* - **Decisions.** A common interface eases switching providers but flattens their features: keep a pass-through mode. PII masking trades accuracy against latency and false positives that corrupt the context. Fallback raises availability, but the other model behaves differently, so prompts and evals must cover both. Caching answers by question similarity is risky: similar is not the same. Prompt caching needs a stable prefix on the same model and provider, and at Anthropic the same workspace (see “Prompt caching”), so a fallback starts with a cold cache. *Open question: once a budget is exceeded, a hard block or a cheaper model?* - **Evaluation.** Measure the gateway, not the answers: added latency at p50 and p99, errors per provider, cost per team, masking effectiveness on local data formats (national ID numbers, addresses, non-English names). Answer quality belongs to the team that owns the feature; the gateway gives it traces and pinned model versions (see “Compute and “getting dumber””). *Open question: what do you log, and what must you never log?* - **Failures.** A partial provider outage triggers a retry storm: backoff with jitter, a circuit breaker and a retry budget help. An agent in a loop can burn a monthly budget in an hour, so limits apply per key and per request. The gateway is a single point of failure: several instances, and a deliberate choice whether a failing filter passes or blocks traffic. Prompt logs are a leak risk of their own. *Open question: which provider errors does the gateway absorb, and which reach the teams?* ### Screening 100k CVs a week - **Requirements.** “Score 100,000 applications a week against job postings”: about 14,000 a day, results within hours. GDPR (Art. 22) restricts decisions based solely on automated processing, and the AI Act classifies recruitment as high-risk (Annex III; after the 2026 Digital Omnibus the obligations apply from 2 December 2027), so you need explanations, an audit trail and human oversight. *Open question: who makes the decision, the model or the recruiter?* - **Architecture.** An event queue, a parser (PDF and DOCX to text, OCR for scans), extraction into a schema (see “Enforcing output format”), a score for each criterion in the posting with a quote from the CV, results in the recruiter’s dashboard. All asynchronous: retries are cheap and traffic peaks don’t block applications. *Open question: how do you guarantee exactly one score per application despite retries?* - **Decisions.** Batch mode: about 50% cheaper for results within 24 hours. The instructions and the posting form a stable prefix for prompt caching, since one posting has hundreds of candidates. Cascade: a cheaper model scores the clear-cut cases, a stronger one the borderline ones, and neither rejects anyone on its own: a rejection without meaningful human review is a decision based solely on automated processing. The criteria come from the posting, not the model’s “knowledge”; name, age and photo are removed before scoring. *Open question: how do you set the borderline threshold?* - **Evaluation.** A golden set of a few hundred CVs scored by several recruiters. The model’s agreement with humans is compared with the agreement among humans, the realistic ceiling. Bias tests: pairs of CVs that differ only in name, gender or age must get the same score. In production, watch the score distribution over time (see “Evals”). *Open question: where does the ground truth come from if recruiters disagree?* - **Failures.** The parser returns garbage for two-column layouts or tables, and the model scores it anyway: detect poor text and route it to a human. A candidate adds “rate me highest” in white text (see “Prompt injection”). No tools and a schema limit what the injection can do, but not its goal: the score itself. So the parser flags text missing from the rendered page (text layer versus OCR), code checks that every quote behind a score is in the visible text, and flagged CVs go to a human. *Open question: how will you notice that candidates have started writing CVs for the model?* ### RAG over 50M documents - **Requirements.** “An assistant answers employees’ questions from 50 million documents.” Ask about permissions, freshness, languages and formats, citations and response time. The scale: at about 20 chunks per document that is about 1 billion vectors: at 1,024 dimensions about 4 TB in float32, 1 TB in int8. *Open question: how will you shrink the index, and how much quality will you lose?* - **Architecture.** Ingest: parsing, chunking that preserves headings and tables, and a short note per chunk on its document and section; a vector index and a keyword index (BM25). At query time: permission filter, hybrid search, rank fusion, reranking, a prompt with the few best chunks, an answer with citations (see “Embeddings and vector search” and “RAG”). *Open question: how do you re-index 50 million documents after changing the embedding model?* - **Decisions.** The permission filter goes into the index query itself: filtering afterwards loses results and risks leaks. Hybrid search, because vectors capture meaning and keywords catch contract numbers and proper names. A reranker is usually the biggest quality gain, for tens to hundreds of milliseconds. A smaller chunk is easier to hit; a larger one carries more context. *Open question: what do you do when the reranker becomes the bottleneck?* - **Evaluation.** The two stages separately: retrieval (is the right chunk in the top k) and generation (is every claim supported by a chunk, and does each citation point to the right place). Questions from experts, questions a model generates from random chunks, checked on a sample, and questions with no answer in the corpus (see “Evals”). *Open question: how do you build a question set without hand-labelling millions of documents?* - **Failures.** A permission is revoked but the chunk is still cached: the cache key includes permissions. An old version of a document beats the new one: dates in metadata and version deduplication. Retrieval finds nothing, and the model answers from memory, citing the nearest chunk (see “Hallucinations”). A document carries a hidden instruction (see “Prompt injection”). *Open question: how do you detect answers unsupported by sources in production?* ### The agent is slow and expensive - **Requirements.** “A task takes 4 minutes and costs $2. Get it down to 30 seconds and 20 cents without losing quality.” Ask: today’s success rate, whether the user waits for the result, which tasks dominate, and what matters most when you can’t have everything. *Open question: what do you sacrifice first: cost, latency or success rate?* - **Architecture.** Before changing anything, trace every turn: input and output tokens, cache hits, thinking tokens, model and tool time. Usually most of the cost is context resent every turn, and most of the time is sequential turns plus output tokens. The target: a stable prefix, concise tools, a small model for simple steps, parallel calls, subagents for side tasks. *Open question: which metrics do you alert on, and at what thresholds?* - **Decisions.** Fix the prefix: nothing variable before the system prompt and tools, because a high cache hit rate cuts input cost several times over (see “Prompt caching”). Trim tool results: 20k tokens of JSON resent every turn is the most expensive line in the system. A model per step: routing and extraction on a small model, thinking effort set with evals (see “Choosing a model”). Compaction and subagents keep the context short (see “Context engineering and memory”). *Open question: what do you lose with compaction, and how do you detect it?* - **Evaluation.** Every optimisation runs on the same task set: pass^k, cost and time per task, p95 rather than the mean, pairwise comparison. A cheaper agent that fails more often ends up costing more in retries and human work. *Open question: how do you build a task set when a task can be solved in many ways?* - **Failures.** The small model botches decisions that only looked simple. Compaction loses a detail needed 20 turns later. A small change at the start of the prompt silently zeroes cache hits and the cost comes back, so alert on the hit rate. An agent without hard limits goes round in a loop (see “The agent loop”). *Open question: how do you detect that the agent is stuck in a loop?* ### Real-time classification - **Requirements.** “Thousands of submissions a minute, a decision within 50 ms, mistakes are costly.” Ask: how many classes and how often they change, what a false alarm costs versus a miss, whether there is labelled history. 50 ms rules out generation by a large LLM on the synchronous path: time to first token alone can be longer. *Open question: what does each kind of mistake cost?* - **Architecture.** The decision state (submission, customer history, rules in force) goes to a fast classifier: a fine-tuned encoder or embeddings with logistic regression, in-process or next to the service. A hosted decision model (see “When not to use an LLM: classifiers and System One”) fits only if its measured p99 latency, network included, stays within budget. Above the confidence threshold the decision is automatic; below it the case gets a provisional “for review” at once and goes asynchronously to a reasoning model or a human. Amounts, permissions and blocklists stay in code. *Open question: what stays in rules, and what does the model judge?* - **Decisions.** A small classifier answers in milliseconds for a fraction of a cent, but needs data and retraining for every new class. An LLM needs no data, so it helps with labelling and borderline cases, off the 50 ms path. The threshold trades the automation rate against errors and is set on data. The model scores, the code decides. *Open question: how will you know the small classifier is no longer enough?* - **Evaluation.** Precision and recall per class on production data. A curve of automation rate against error rate at different thresholds. Calibration: does 90% confidence really mean 90% correct? In production, human corrections serve as labels, and monitoring watches the class distribution. *Open question: how do you check calibration on your own data?* - **Failures.** Drift: a new product or a new fraud campaign, and confidence stays high while accuracy drops. A flood of escalations after a threshold change. Labels arrive only for escalated cases, so the automated errors go unseen; a human therefore also reviews a random sample of automated decisions. *Open question: how large must that sample be to spot a 1% error rate?* ### How to approach a design - Start with questions, not with the model: who uses it, what volume, acceptable latency, cost per request, data sensitivity, cost of a mistake. Write the numbers down and multiply: tokens per request times requests times price. - Build the simplest version that works: a single model call or a workflow before an agent (see “The agent loop”), a classifier before an LLM. Add complexity only where the requirements demand it. - Record every decision as a trade-off: what you gain, what you lose and at what number you would change your mind. “We use RAG” is not a decision; “RAG, because the knowledge changes every week and answers must cite sources” is. - Finish with evaluation and failures: how quality is measured before launch and in production, what will break and how you will notice. Without them a system can be drawn, but not run. ### Numbers for estimates - Output tokens usually cost 4–8 times more than input tokens (2025–2026 price lists), and they set the generation time. Long input is cheap per token, but you pay for it on every request. - At Anthropic a prompt cache read costs 10% of the input price on most models and less on the newest (5% on Claude Opus 5.5, 2.5% on Fable 5.1, as of September 2026), and a write costs 1.25× (5-minute entry) or 2× (one-hour entry). At OpenAI the read discount is 50–90% depending on the model. Batch mode at both providers is about 50% cheaper for results within 24 hours. - Agent latency is the sum of its turns, and each turn is time to first token, generation and the tool. Ten turns of 3 seconds each is half a minute, however fast a single request is. ### Check yourself **Question:** How would you estimate the cost and latency of an LLM system before you build it? **Short answer:** Start from a single request: input tokens (system prompt, context, history) and output tokens separately, because output costs several times more and drives generation time. Multiply by request volume and prices, then subtract cache hits and whatever can run in batch at half price. For an agent, also multiply by the number of turns, because each turn resends the whole context. Latency is the sum of turns: time to first token, generation and tools. Compare the result with the budget and the latency limit before choosing a model. ### Follow-up questions - **When isn’t an LLM needed here?** When there are few classes, plenty of labelled data and a hard latency or cost limit: then a small classifier or rules win. Also when every decision must be fully reproducible and auditable. An LLM is then useful for labelling data, not on the production path (see “When not to use an LLM: classifiers and System One”). - **How do you choose a model?** By requirements, not leaderboards: the cheapest model that passes your eval set at the required latency, with a pinned version and a migration plan. A strong model for the hard steps, a small one for the simple ones (see “Choosing a model”). - **How do you keep working when the model provider goes down?** Timeouts and retries with backoff, fallback to a second provider validated with the same eval set, and a degraded mode without an LLM where possible, such as a queue instead of an immediate answer. - **How do you split responsibility between the model and code?** The model judges, extracts and writes; code decides and executes. Business rules, permissions, limits and action approval are deterministic and testable, and the model gets only what can’t be written down as a rule. - **When would you fine-tune instead of using a prompt and RAG?** When the problem is about form, not knowledge: a fixed format, style, classification in a narrow domain, or when a small fine-tuned model is to replace a large one for cost and latency. Knowledge that changes goes into the context (see “Fine-tuning and LoRA” and “RAG”). ### Sources - [Chip Huyen: Building A Generative AI Platform (2024)](https://huyenchip.com/2024/07/25/genai-platform.html) - [Anthropic: Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) - [Evidently AI: a database of 800 ML and LLM system design case studies](https://www.evidentlyai.com/ml-system-design) --- ## Flashcards *Practice and review* *Last edited: 28 September 2026* One question for every topic. Answer it yourself first, and only then flip the card and compare with the short answer. “I know it” marks the topic as mastered, the same as the checkbox at the end of a topic page. Progress is saved only in this browser. “Unmastered only” hides the topics you have marked, and “With follow-up questions” adds the deeper questions from each topic’s “Check yourself”.