Home / Chapter 2 · Where a model’s knowledge comes from
Last edited · 7 min read
Reasoning models
A reasoning model writes a chain of thought before it answers: it tries, checks and corrects itself. It is the same token-by-token machinery, trained with RL on tasks with verifiable outcomes, and you pay for thinking tokens as output.
In plain wordsA student allowed to work on scratch paper during a test makes fewer mistakes than one who writes the answer straight down. But scratch paper won’t help them recall a date they never learned. Working it out takes time, and for a model every line of it goes on the bill.
Pick a task and a thinking level. See where thinking improves the result and where it only adds cost
Chance of a correct answer
Output tokens including the answer
Illustrative numbers, not a benchmark result or a measurement. Price $10 per million output tokens, 80 tokens per second, output only, excluding input and time to first token. The chain of thought shown is abridged; a real one is often many times longer. The answer is a single example sample. Not every model lets you switch thinking off.
Where thinking comes from
- There is no separate thinking module. The chain of thought is ordinary tokens from the same loop as the answer (see “The next token”), just marked as thinking. Every step written down stays in the context, so later tokens build on it like on scratch paper.
- The model learns to write this scratch paper through RL (see “Training”) on automatically checked tasks: maths with a known answer, code with tests. The reward goes mainly to the result, so whatever leads to it gets reinforced, including self-checking and backtracking. In DeepSeek-R1-Zero these behaviours became frequent through RL alone, without human-written examples, though base models already show them occasionally (Liu et al., 2025). RLVR mostly makes the model reach, more reliably, answers the base model could already find: given enough samples (large pass@k), the base model catches up (Yue et al., 2025).
- It is one form of test-time compute: spending more computation at answer time instead of building a bigger model. Others are sampling several attempts and voting, or picking the best one with a verifier. Gains shrink as thinking gets longer, and on some tasks longer thinking makes results worse (inverse scaling in test-time compute, 2025).
When it pays off
- It helps on multi-step tasks where mistakes come easily: calculations, code, planning, analysis with many conditions. On simple facts, lookups and short classification it only adds cost and waiting.
- It adds no knowledge. What the model doesn’t know it won’t work out: you get a longer, more confident-sounding guess (see “Hallucinations”).
- Every thinking token is a decode step that waits on memory (see “Why the GPU is idle”). A few thousand thinking tokens means tens of seconds before the first word of the answer. In a UI, show progress or a summary of the thinking. Run tasks with no human in the loop in the background or as batch jobs.
Controls and the bill
- You control the effort level (
reasoning.effortin OpenAI,effortin Anthropic,thinking_levelin Gemini, fromnonetomaxdepending on the model). A fixed thinking-token budget (budget_tokens,thinking_budget) survives only on older models: Claude 4.7 and later reject it with a 400, and Gemini 3 accepts it only for backward compatibility. Even there it is a target, not an allocation. You pick the level with evals on your own tasks (see “Evals”). - On the newest flagship models thinking can’t be switched off (as of September 2026): Claude Opus 5.5 and Fable 5.1 reject
thinking: {type: "disabled"}with a 400, GPT-6 Astra rejectsnonethe same way, and in Gemini 3.x the lowest level isminimalorlow. You choose how much to think, not whether. In adaptive mode the model decides per request, and at lower effort it skips thinking on simple requests more often. - Thinking tokens are billed as output, even when you can’t see them. APIs usually don’t return the raw chain of thought, only a summary or an empty block with encrypted content, so
usageshows more tokens than you can see in the response. - Thinking counts towards the output limit (
max_tokens,max_output_tokens) and the context window. Set the limit too low and the response is cut off mid-thought. OpenAI recommends reserving at least 25,000 tokens for reasoning and output when you start. - The sampler is usually locked. Claude Opus from 4.7 rejects any non-default
temperature,top_portop_kwith a 400, with or without thinking, and OpenAI with reasoning enabled doesn’t accepttemperatureortop_p(see “The next token”).
What to watch out for
- The chain of thought need not faithfully describe how the answer came about. In an Anthropic study (2025), models used a planted hint but admitted it only in a minority of cases: Claude 3.7 Sonnet 25% of the time, DeepSeek R1 39%. Read it as a debugging clue, not an audit.
- In an agent the model also thinks between tool calls (interleaved thinking, see “The agent loop”). You send thinking blocks back unchanged in the next request, with their signature or encrypted content. The API rejects a modified block, and an omitted one breaks the continuity of reasoning.
- On Claude Opus 5.5 and Fable 5.1 a block is also bound to everything before it. If your own code changes the system prompt, the tools or an earlier message (for example by trimming an old tool result), replaying the block returns a 400 on accounts created from 31 August 2026 (older ones only when they opt in). Server-side context editing and compaction don’t count as changes. Keep the history append-only and add new instructions as new messages.
- Some models strip thinking from previous turns, others keep it. Kept thinking takes up the window and is billed as input. Stripped thinking changes the prefix from that point, so the cache misses after it. Changing the effort level or budget mid-conversation also invalidates the prompt cache, because the setting goes into the prompt (see “Prompt caching”).
Check yourself
What are reasoning models, and how much thinking is worth paying for?
A reasoning model generates a chain of thought before answering: it lays out steps, checks them and fixes mistakes. It is the same next-token machinery, trained with reinforcement learning on tasks with verifiable outcomes such as maths or code with tests. Thinking tokens are billed as output, count towards limits and add latency. It helps on multi-step problems; on simple facts and classification it only adds cost, and it adds no knowledge. The newest models don’t let you switch thinking off, so pick the effort level with evals. The visible reasoning is not a faithful explanation of the answer.
Po polsku
Model rozumujący przed odpowiedzią generuje tok myślenia: rozpisuje kroki, sprawdza je i poprawia błędy. To ten sam mechanizm przewidywania następnego tokena, wyuczony w RL na zadaniach ze sprawdzalnym wynikiem, jak matematyka czy kod z testami. Tokeny myślenia płaci się jak wyjście, liczą się do limitów i wydłużają czekanie. Pomaga przy zadaniach wieloetapowych, przy prostych faktach i klasyfikacji tylko kosztuje, a wiedzy nie dodaje. W najnowszych modelach myślenia nie da się wyłączyć, więc poziom wysiłku dobiera się ewaluacją. Widoczny tok myślenia nie jest wiernym wyjaśnieniem odpowiedzi.
Follow-up questions (4)
- How is a reasoning model different from a “think step by step” prompt?
- A chain-of-thought prompt asks an ordinary model to lay out its steps, and that helps too. A reasoning model was trained to do it with RL: it writes much longer chains of thought, checks and corrects itself more often, and the API keeps the thinking separate from the answer.
- Why not always set the highest level?
- Thinking tokens cost as much as output and each one adds latency, while the quality gain shrinks. On simple tasks the gain is zero, and sometimes longer thinking makes the result worse. Pick the level per task type on your own eval set.
- Does the chain of thought explain why the model answered the way it did?
- There is no such guarantee. Research shows that models can use a hint without mentioning it in their chain of thought. On top of that, the API often returns only a summary. It is useful for debugging a prompt, not as an audit.
- How does thinking work in an agent with tools?
- The model also thinks between tool calls, after each result. You send thinking blocks back unchanged in the next request, with their signature or encrypted content, so the reasoning stays continuous and the cache hits. On Claude Opus 5.5 and Fable 5.1, changing anything before a block ends in an error, so an agent’s history is append-only.