## Hallucinations

*Quality and security*

*Last edited: 28 September 2026*

A model has no “I don’t know” mode: it always writes a plausible continuation, in the same confident tone whether it knows the answer or is guessing. Hallucinations cannot be switched off; they can be reduced and measured.

**In plain words:** A student in an oral exam who never says “I don’t know”. When they know the answer, they speak fluently and confidently. When they don’t, they speak just as fluently and confidently, so you can’t tell the difference from the tone.

*Interactive widget on the page: Turn on a source in the context and permission to say “I don’t know”, separately and together. Watch which fabrications disappear, which remain and how many answers you lose along the way.*

### Where they come from

- The model predicts the next token and always returns one. There is no separate signal for “my knowledge ends here”, and a guessed sentence is as fluent as a true one. A sampled token cannot be taken back: if the answer began with “In 1978”, the rest of the sentence will justify that year (see “The next token”).
- Facts that appeared once or never in the training data are stored in the weights weakly or not at all. The model saw the capital of Australia thousands of times, and the opening year of a small museum once at most. Kalai et al. (OpenAI, 2025) show that for facts with no learnable pattern, such as birthdays, the hallucination rate after pretraining is roughly at least the fraction of facts that appeared exactly once in the data.
- Post-training does not remove this, because almost all popular benchmarks grade 0/1: “I don’t know” scores zero, the same as a wrong answer. Under such grading abstaining is never optimal, so a model trained and selected for the score learns to guess, like a student on a test with no negative marking.

### Types

- Fabricated fact: a wrong date, number or name stated confidently. Fabricated source: a paper title, court ruling, link or quote that does not exist. In 2023 a federal court in New York fined lawyers $5,000 for a filing that cited rulings invented by ChatGPT (Mata v. Avianca).
- Unfaithfulness to the source: the document is in the context, yet the answer contradicts it or adds something it does not contain. In RAG this is the typical generation-stage error, and a dangerous one, because the citation makes the answer look verified. Reasoning error: the facts are right, but the arithmetic is wrong or the conclusion does not follow from the premises.
- In an agent: a tool call with an invented argument (customer ID, file path, field name), or a report such as “I sent the email” or “the tests pass” when no tool did any such thing. In code: functions, parameters and whole packages that do not exist. An attacker can register a package under a name models tend to invent and wait for an agent to install it.

### What helps, and its limits

- A source in the context with citations (see “RAG”) helps most, but only when retrieval finds the passage containing the answer. Add explicit permission to say “I don’t know” and the instruction “first extract verbatim quotes from the source, then answer based only on them”. A quote can be checked with a plain text search. The price is that the model answers fewer questions.
- Tools instead of memory: code or a calculator for arithmetic, search and a database for facts, documentation for APIs (see “Tools (function calling)”). A verification step, meaning a separate call that reads the answer together with the sources and strikes out unsupported claims, catches some of the errors but makes mistakes of its own.
- Lower temperature is not a cure: temperature 0 picks the most probable token, so a fabricated date simply becomes reproducible. Reasoning before answering (see “Reasoning models”) helps with logic and arithmetic, because the model can check its steps, but it cannot add knowledge that is not in the weights. Fine-tuning is a poor way to add it: the model learns new facts slowly, and as it learns them its tendency to fabricate grows (Gekhman et al., 2024).

### Confidence threshold and measurement

*Interactive widget on the page: Move the confidence threshold below which the model says “I don’t know”. Compare the score under 0/1 grading and under grading with a penalty for errors, and find the point where abstaining pays off.*

- Choose the threshold by the cost of a mistake and measure it on your own set (see “Evals”): correct, wrong and abstained answers separately, plus questions whose answer is deliberately missing from the sources. Accuracy alone rewards guessing; error rate alone rewards silence.
- Model confidence from `logprobs` is not a measure of truth. A model can be confident in an error, confidence is spread across different phrasings of the same answer, and post-training hurts calibration: in the GPT-4 report the base model was well calibrated on MMLU questions, and noticeably worse after post-training. Without sources, consistency is a better signal: sample the same answer several times and check whether the samples mean the same thing (semantic entropy, Farquhar et al., Nature 2024). Guessed facts change between samples; known ones stay put.
- Check faithfulness and citations in code. First, that the cited passage exists verbatim in the document. Native citations in APIs, for example in Claude, guarantee a correct pointer into the document, but not that the passage supports the claim. Then, claim by claim, whether the passage supports it: an NLI model or an LLM judge with a yes/no verdict.

### Check yourself

**Question:** Where do hallucinations come from, and how do you reduce and measure them in production?

**Short answer:** The model always produces a plausible continuation and has no built-in “I don’t know”, so on rare facts it guesses in the same confident tone. Training and benchmarks graded 0/1 reward a hit and give zero for abstaining, so guessing pays. It can’t be switched off, only reduced: sourced context with citations, permission to say “I don’t know”, tools for facts and arithmetic, and checking citations in code. Lower temperature just repeats the same guess. Measure errors, abstentions and faithfulness to sources separately, and set the answer threshold by the cost of a mistake.

### Follow-up questions

- **Can you detect a hallucination from logprobs?** Partly. Low confidence can be a signal, but a model can be confident in an error, confidence is spread across different phrasings of the same answer, and post-training hurts calibration. A better signal comes from sampling several answers and checking whether they mean the same thing.
- **Why does a RAG system still make things up?** Retrieval always returns something. If the chunk doesn’t contain the answer, the model often fills the gap from its weights and attaches a citation to the nearest source. On top of that, it can distort the chunk right in front of it.
- **The agent reports “tests pass”, but they don’t. What do you do?** Trust the state, not the report. Code runs the tests and sends the result back to the model, and a tool, not the model’s claim, checks the task’s completion condition. In evals you grade the end state, and runs with a false report go into the set as new cases.
- **How do you measure hallucinations in your own system?** A set of questions with verified answers plus questions whose answer isn’t in the sources. Count correct, wrong and abstained answers separately, and for RAG also faithfulness: whether every claim is supported by the supplied chunk.

### Sources

- [Kalai et al.: Why Language Models Hallucinate (arXiv, 2025)](https://arxiv.org/abs/2509.04664)
- [Claude docs: Reduce hallucinations](https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-hallucinations)
- [Lilian Weng: Extrinsic Hallucinations in LLMs](https://lilianweng.github.io/posts/2024-07-07-hallucination/)

Interactive page: https://howaiworks.dev/hallucinations/
