Home / Chapter 2 · Where a model’s knowledge comes from
    Last edited · 10 min read

    Use with AI

    Distillation

    Distillation trains a small model, the student, on the outputs of a large one, the teacher. It is why small models beat what their size suggests and how reasoning reached 7B models, and it lets you replace an expensive model on one task with a cheaper one.

    In plain wordsA chess student who sees only a grandmaster’s moves learns more slowly than one who is also told which other moves were nearly as good and which lose at once. On-policy distillation is the grandmaster watching the student’s own games and marking every move, instead of showing their own. The student rarely surpasses the master, but does pick up the master’s mistakes.

    Switch the target the student learns from and change the temperature. Compare what one label says with what the teacher’s distribution says

    bits of entropy in the target: 0 is one answer, 3 is eight equal ones
    tokens with at least 5% of the target

    Real data: the eight most likely next tokens from Qwen3-0.6B-Base, rescaled to 100%. This small model plays the teacher here; a real teacher is larger, and its distribution covers the whole vocabulary. The hard label is the token the teacher would write at temperature 0. The 5% threshold is this widget’s choice.

    Hard labels and distributions

    On-policy distillation

    In the pretraining of small models

    Reasoning distillation

    Limits

    The labs’ side

    Distillation in your system

    Check yourself

    How does distillation move a large model’s ability into a small one, and how do hard labels, teacher distributions and on-policy distillation differ?

    The student, a small model, learns to imitate the teacher, a large one. Most often with hard labels: the teacher generates answers and the student is trained on them with ordinary SFT, so text from an API is enough. Logit distillation minimises the KL divergence between the teacher’s and the student’s full next-token distributions, usually with a temperature. Each example then carries more information, because it also shows which alternatives are close, but it needs the teacher’s logits and a shared tokeniser. In on-policy distillation the student generates and the teacher grades each of its tokens with reverse KL, so the student also learns from its own mistakes; in Qwen3 this beat RL at about a tenth of the GPU hours. The student won’t surpass the teacher, copies its errors and picks up style more easily than skill, so test it on your own eval set. The terms of OpenAI, Anthropic and Google forbid using outputs to train competing models.

    Po polsku

    Uczeń, mały model, uczy się naśladować nauczyciela, duży model. Najczęściej na etykietach twardych: nauczyciel generuje odpowiedzi, a uczeń przechodzi na nich zwykłe SFT, więc wystarczy tekst z API. Destylacja logitów minimalizuje dywergencję KL między pełnym rozkładem następnego tokena nauczyciela i ucznia, zwykle z temperaturą. Jeden przykład niesie wtedy więcej informacji, bo pokazuje też, które alternatywy są bliskie, ale potrzebne są logity nauczyciela i wspólny tokenizer. W destylacji on-policy uczeń generuje sam, a nauczyciel ocenia każdy jego token odwrotną KL, więc uczeń uczy się też na własnych błędach; w Qwen3 dało to lepszy wynik niż RL przy ok. 1/10 godzin GPU. Uczeń nie przerośnie nauczyciela, kopiuje jego błędy i łatwiej przejmuje styl niż umiejętności, więc sprawdzaj go na własnym zestawie ewaluacyjnym. Regulaminy OpenAI, Anthropic i Google zabraniają używania wyników do trenowania modeli konkurencyjnych.

    Follow-up questions (5)
    If distributions carry more information, why is most distillation done with hard labels?
    Because hard labels need only text. A distribution needs the teacher’s logits, so open weights, and an API returns at most a short list of logprobs, or none. Both models need the same tokeniser, and a full distribution is 100–260k numbers per position, so a sample of tokens is stored instead, such as Gemma 3’s 256. With a teacher behind an API, SFT on its answers is what remains.
    Why does on-policy distillation beat SFT on the same teacher’s answers?
    SFT trains in situations the teacher reaches. When generating, the student makes its own mistakes, lands in situations it never saw, and the errors compound. On-policy distillation trains on the student’s own text while the teacher grades every token, so the signal is dense like SFT and comes from the student’s distribution like RL. The price is a teacher forward pass on every student sample and access to its logprobs.
    The student matches the teacher on the eval set, but in production it fails on slightly different inputs. What happened?
    The distillation data covered the eval set’s distribution rather than production traffic, and the student picked up the teacher’s style more easily than its skill. What helps: broadening the inputs with real production cases, including hard and unusual ones, a fresh eval set drawn from production, and routing to the teacher the cases where a validator rejects the student’s result.
    May you distil a model available only through an API into your own model?
    Technically, hard labels are enough. The terms (as of September 2026) forbid using outputs to train competing models. OpenAI makes exceptions for undistributed classifiers and embeddings and for fine-tuning its own models, and Anthropic allows it only with express approval. The terms don’t say whether a narrow model for your own use “competes”, so that is a question for a lawyer. An open teacher whose licence allows distillation, such as DeepSeek-R1 under MIT, avoids the question.
    Does a bigger teacher always make a better student?
    No. In Gemma 3’s ablation the smaller teacher won with short training and the larger one only with long training. Models up to 3B don’t consistently gain from strong teachers’ long reasoning traces, and a mix of short and long traces helps them (Li et al., 2025). Pick the teacher by the student’s score on the eval, not by the teacher’s own score.

    Sources

    Report an error · Suggest a fix