Home / Chapter 2 · Where a model’s knowledge comes from
    Last edited · 7 min read

    Use with AI

    Training

    A model is built in three stages. Pretraining provides language and knowledge through next-token prediction, SFT teaches the assistant role, and RL polishes behaviour and reasoning. After training the weights are frozen, so anything the model doesn’t know has to be given to it in context.

    In plain wordsFirst, years of reading everything in sight. Then a vocational course on example conversations, after which it knows how an assistant behaves. Finally, graded practice: it tries on its own and an examiner rewards good results, so it learns whatever earns points.

    Click a stage and compare what the model does with the same prompt

    Answers written by hand, for illustration.

    Pretraining

    Post-training: SFT and RL

    What this means in practice

    Check yourself

    How is an LLM made: what is the difference between pretraining, SFT, RLHF and RLVR?

    Pretraining teaches next-token prediction on trillions of tokens of text with a cross-entropy loss. That gives language and knowledge, but the base model only continues documents. SFT on conversations in the chat template teaches the assistant role and the tool format. RLHF optimises against a reward model learned from comparisons made by people or a model judge, with a KL penalty for drifting away from the SFT model. RLVR optimises against a reward checked by a program, such as tests: that is how reasoning models are made, and, in multi-turn tool environments, models for agents. The risk of RL is reward hacking. After training the weights are frozen, so current knowledge has to be supplied in context.

    Po polsku

    Pretraining uczy przewidywać następny token na bilionach tokenów tekstu, stratą cross-entropy. Stąd język i wiedza, ale model bazowy tylko dokańcza dokumenty. SFT na rozmowach w szablonie czatu uczy roli asystenta i formatu narzędzi. RLHF optymalizuje pod model nagrody wyuczony z porównań ludzi albo modelu-sędziego, z karą KL za odejście od modelu po SFT. RLVR optymalizuje pod nagrodę sprawdzaną programem, np. testami: tak powstają modele rozumujące, a w wieloturowych środowiskach z narzędziami także modele do agentów. Ryzykiem RL jest reward hacking. Po treningu wagi są zamrożone, więc bieżącą wiedzę trzeba podać w kontekście.

    Follow-up questions (4)
    RLHF versus RLVR?
    RLHF takes its reward from a model trained on the preferences of people or a model judge: good for style and helpfulness, but the reward model can be fooled. RLVR takes its reward from an automatic check, such as tests or the task’s result: harder to fool and easy to scale, but only where a program can verify the result.
    PPO, DPO, GRPO?
    PPO is classic RL with a reward model and a separate critic that estimates the value of a state. DPO learns directly from “better/worse answer” pairs, with no reward model and no generation in the loop. GRPO (DeepSeek) generates several answers to the same prompt and compares them with each other, so it needs no critic.
    Why a KL penalty in RL?
    It keeps the model close to the SFT model. Without it, optimisation quickly finds answers that the reward model rates highly and people rate poorly, and the model loses general skills. RLVR for reasoning often drops it (DAPO), because the model has to move far from its starting point; there a good verifier is what stops the model gaming the reward.
    When fine-tuning and when RAG?
    Fine-tuning for style, format and narrow skills. RAG for knowledge that changes and needs a source (see “Fine-tuning and LoRA” and “RAG”).

    Sources

    Report an error · Suggest a fix