Home / Chapter 2 · Where a model’s knowledge comes from
    Last edited · 7 min read

    Use with AI

    Reasoning models

    A reasoning model writes a chain of thought before it answers: it tries, checks and corrects itself. It is the same token-by-token machinery, trained with RL on tasks with verifiable outcomes, and you pay for thinking tokens as output.

    In plain wordsA student allowed to work on scratch paper during a test makes fewer mistakes than one who writes the answer straight down. But scratch paper won’t help them recall a date they never learned. Working it out takes time, and for a model every line of it goes on the bill.

    Pick a task and a thinking level. See where thinking improves the result and where it only adds cost

    Task:
    Thinking:
    thinking tokens
    generation time
    cost of 1000 such requests

    Chance of a correct answer

    Output tokens including the answer

    Illustrative numbers, not a benchmark result or a measurement. Price $10 per million output tokens, 80 tokens per second, output only, excluding input and time to first token. The chain of thought shown is abridged; a real one is often many times longer. The answer is a single example sample. Not every model lets you switch thinking off.

    Where thinking comes from

    When it pays off

    Controls and the bill

    What to watch out for

    Check yourself

    What are reasoning models, and how much thinking is worth paying for?

    A reasoning model generates a chain of thought before answering: it lays out steps, checks them and fixes mistakes. It is the same next-token machinery, trained with reinforcement learning on tasks with verifiable outcomes such as maths or code with tests. Thinking tokens are billed as output, count towards limits and add latency. It helps on multi-step problems; on simple facts and classification it only adds cost, and it adds no knowledge. The newest models don’t let you switch thinking off, so pick the effort level with evals. The visible reasoning is not a faithful explanation of the answer.

    Po polsku

    Model rozumujący przed odpowiedzią generuje tok myślenia: rozpisuje kroki, sprawdza je i poprawia błędy. To ten sam mechanizm przewidywania następnego tokena, wyuczony w RL na zadaniach ze sprawdzalnym wynikiem, jak matematyka czy kod z testami. Tokeny myślenia płaci się jak wyjście, liczą się do limitów i wydłużają czekanie. Pomaga przy zadaniach wieloetapowych, przy prostych faktach i klasyfikacji tylko kosztuje, a wiedzy nie dodaje. W najnowszych modelach myślenia nie da się wyłączyć, więc poziom wysiłku dobiera się ewaluacją. Widoczny tok myślenia nie jest wiernym wyjaśnieniem odpowiedzi.

    Follow-up questions (4)
    How is a reasoning model different from a “think step by step” prompt?
    A chain-of-thought prompt asks an ordinary model to lay out its steps, and that helps too. A reasoning model was trained to do it with RL: it writes much longer chains of thought, checks and corrects itself more often, and the API keeps the thinking separate from the answer.
    Why not always set the highest level?
    Thinking tokens cost as much as output and each one adds latency, while the quality gain shrinks. On simple tasks the gain is zero, and sometimes longer thinking makes the result worse. Pick the level per task type on your own eval set.
    Does the chain of thought explain why the model answered the way it did?
    There is no such guarantee. Research shows that models can use a hint without mentioning it in their chain of thought. On top of that, the API often returns only a summary. It is useful for debugging a prompt, not as an audit.
    How does thinking work in an agent with tools?
    The model also thinks between tool calls, after each result. You send thinking blocks back unchanged in the next request, with their signature or encrypted content, so the reasoning stays continuous and the cache hits. On Claude Opus 5.5 and Fable 5.1, changing anything before a block ends in an error, so an agent’s history is append-only.

    Sources

    Report an error · Suggest a fix