## Distillation

*Where a model’s knowledge comes from*

*Last edited: 28 September 2026*

Distillation trains a small model, the student, on the outputs of a large one, the teacher. It is why small models beat what their size suggests and how reasoning reached 7B models, and it lets you replace an expensive model on one task with a cheaper one.

**In plain words:** A chess student who sees only a grandmaster’s moves learns more slowly than one who is also told which other moves were nearly as good and which lose at once. On-policy distillation is the grandmaster watching the student’s own games and marking every move, instead of showing their own. The student rarely surpasses the master, but does pick up the master’s mistakes.

*Interactive widget on the page: Switch the target the student learns from and change the temperature. Compare what one label says with what the teacher’s distribution says.*

### Hard labels and distributions

- Hard labels (sequence-level distillation): the teacher generates answers to your prompts, and the student is trained on them with ordinary SFT (see “Training”). Text is all you need, so it works through an API and across different tokenisers. This is the most common form of distillation.
- Distributions (logit distillation): at every position the student minimises the KL divergence from the teacher’s full next-token distribution to its own. Hinton et al. (2015) raise the softmax temperature in both models (see “The next token”) to bring out small probabilities, and multiply this part of the loss by T², because its gradients shrink as 1/T². Which wrong answer is almost right lives in the ratios of very small probabilities; Hinton called this “dark knowledge”.
- A distribution carries more information per example: a ranking of alternatives instead of a single index, and a gradient that varies less between examples. In the paper’s speech experiment, on 3% of the data hard labels overfitted at 44.5% accuracy, while the teacher’s distributions reached 57.0%, against 58.9% for training on all the data.
- The price: you need the teacher’s logits, so open weights or your own model. An API returns at most a short list of logprobs, and Claude, or OpenAI with reasoning on, none at all. Both models need the same tokeniser, because the distribution is over token IDs. A full distribution is 100–260k numbers per position, so Gemma 3 stores 256 tokens sampled by teacher probability, and zeroes and renormalises the rest.

### On-policy distillation

- Both methods above train on the teacher’s text. When generating, the student makes mistakes the teacher never makes, lands in situations it never saw in training, and the errors compound.
- In on-policy distillation the student generates the answer and the teacher, in one forward pass, computes the probability of each of its tokens. The loss is a per-token reverse KL: the less likely a token is for the teacher, the bigger the penalty. It works like RL, but with every token graded instead of one reward for the whole answer. Reverse KL pushes the student towards one of the teacher’s ways of answering instead of spreading across several.
- Qwen3 (2025) trained its 0.6B to 14B and 30B-A3B models this way, with Qwen3-32B or 235B-A22B as the teacher: first SFT on the teacher’s answers, then an on-policy stage. On Qwen3-8B that stage reached 74.4% on AIME’24 in 1,800 GPU hours, while RL from the same starting point reached 67.6% in 17,920 hours. Thinking Machines spelled out the method in “On-Policy Distillation” (October 2025).

### In the pretraining of small models

- Gemma 2 (2024) trained its 2B and 9B models on a larger teacher’s distributions, on more than 50 times the compute-optimal number of tokens. In an ablation, a 2B model after 500B tokens averaged 67.7 points with distillation from a 7B model and 60.3 without it. In Gemma 3 (2025) every size is distilled.
- Llama 3.2 1B and 3B (September 2024) were pruned from Llama 3.1 8B, and during pretraining the logits of the 8B and 70B models served as targets for every token.
- The usual pretraining target is the one token a document happened to contain; a teacher’s distribution says what could have been there, so every token teaches more. That is one reason 1–4B models beat their size.

### Reasoning distillation

- DeepSeek-R1 (January 2025): about 800k examples, including about 600k reasoning traces kept only when their result was correct. SFT alone on them, with no RL, produced models from Qwen2.5 1.5B to Llama 3.3 70B. Qwen2.5-32B scored 72.6% on AIME 2024 after distillation, while the same base after more than 10k steps of large-scale RL scored 47.0%.
- A small, carefully chosen set also works. s1 (2025): 1,000 questions with reasoning traces from Gemini Flash Thinking, and 26 minutes of training Qwen2.5-32B on 16 H100s. LIMO (2025): 800 selected solutions and 63.3% on AIME24. The knowledge is already in the base; the traces teach the form of long reasoning (see “Reasoning models”).

### Limits

- The student is capped by the teacher and copies its errors and biases; the R1 authors note that going further needs a stronger base and more RL. Hard labels even lock in guesses: if the teacher sampled a name it didn’t know, the student learns to state it with confidence.
- Too big a gap between teacher and student also hurts. In Gemma 3’s ablation the smaller teacher won with short training and the larger one only with long training. Li et al. (2025): models up to 3B don’t consistently gain from strong teachers’ long reasoning traces and learn better from a mix of short and long ones.
- Style transfers more easily than skill. Gudibande et al. (2023): raters judged models trained on ChatGPT’s answers competitive, but targeted tests showed they closed almost none of the gap outside tasks well covered in the data. Scores on benchmarks close to the distillation data transfer; robustness to unusual inputs, much less.

### The labs’ side

- Terms of service forbid using outputs to train competing models (as of September 2026). OpenAI Terms of Use (effective 1 January 2026): “Use Output to develop models that compete with OpenAI.” The API’s Services Agreement (same date) makes two exceptions: undistributed models that categorise, classify or organise data, such as embeddings and classifiers, and fine-tuning OpenAI’s own models. Anthropic Commercial Terms (effective 17 June 2025): customers may not “access the Services to build a competing product or service, including to train competing AI models” unless Anthropic expressly approves. Google Gemini API terms (effective 23 March 2026): “You may not use the Services to develop models that compete with the Services.”
- The raw chain of thought is hidden. For o1 (September 2024) OpenAI gave user experience, competitive advantage and the option to monitor the chain of thought as its reasons. Claude’s documentation says summarising “prevents misuse” and that no setting returns the raw chain of thought. Gemini returns summaries without a stated reason.
- Anthropic (23 February 2026) says it detected campaigns by DeepSeek, Moonshot AI and MiniMax: over 16 million exchanges through about 24,000 fraudulent accounts, aimed at reasoning, tool use and coding, including requests to write out, step by step, the reasoning behind a finished answer. These are one party’s claims.

### Distillation in your system

- The typical case: a large model with a long prompt does one task well but is too expensive or too slow at your volume. Collect a few thousand real inputs, generate answers with the teacher, discard bad ones with a validator or a judge, and fine-tune a small model, for example with LoRA (see “Fine-tuning and LoRA”). The student no longer needs the long prompt, so you pay for fewer input tokens and get answers faster. Compare student candidates as in “Choosing a model”.
- Distributions and on-policy distillation become options when the teacher has open weights and the same tokeniser, such as a larger model from the same family. OpenAI explicitly allows fine-tuning its own models and building internal classifiers, but none of these terms says whether a narrow generative model for your own use “competes”, so that is a question for a lawyer. An open teacher with a suitable licence avoids it: DeepSeek-R1’s MIT licence names distillation explicitly.
- Test the student on an eval set of real cases, set aside before you generate the data (see “Evals”): side by side with the teacher, with unusual cases and refusals included. Cases where the student loses can be routed to the teacher.

### Check yourself

**Question:** How does distillation move a large model’s ability into a small one, and how do hard labels, teacher distributions and on-policy distillation differ?

**Short answer:** The student, a small model, learns to imitate the teacher, a large one. Most often with hard labels: the teacher generates answers and the student is trained on them with ordinary SFT, so text from an API is enough. Logit distillation minimises the KL divergence between the teacher’s and the student’s full next-token distributions, usually with a temperature. Each example then carries more information, because it also shows which alternatives are close, but it needs the teacher’s logits and a shared tokeniser. In on-policy distillation the student generates and the teacher grades each of its tokens with reverse KL, so the student also learns from its own mistakes; in Qwen3 this beat RL at about a tenth of the GPU hours. The student won’t surpass the teacher, copies its errors and picks up style more easily than skill, so test it on your own eval set. The terms of OpenAI, Anthropic and Google forbid using outputs to train competing models.

### Follow-up questions

- **If distributions carry more information, why is most distillation done with hard labels?** Because hard labels need only text. A distribution needs the teacher’s logits, so open weights, and an API returns at most a short list of logprobs, or none. Both models need the same tokeniser, and a full distribution is 100–260k numbers per position, so a sample of tokens is stored instead, such as Gemma 3’s 256. With a teacher behind an API, SFT on its answers is what remains.
- **Why does on-policy distillation beat SFT on the same teacher’s answers?** SFT trains in situations the teacher reaches. When generating, the student makes its own mistakes, lands in situations it never saw, and the errors compound. On-policy distillation trains on the student’s own text while the teacher grades every token, so the signal is dense like SFT and comes from the student’s distribution like RL. The price is a teacher forward pass on every student sample and access to its logprobs.
- **The student matches the teacher on the eval set, but in production it fails on slightly different inputs. What happened?** The distillation data covered the eval set’s distribution rather than production traffic, and the student picked up the teacher’s style more easily than its skill. What helps: broadening the inputs with real production cases, including hard and unusual ones, a fresh eval set drawn from production, and routing to the teacher the cases where a validator rejects the student’s result.
- **May you distil a model available only through an API into your own model?** Technically, hard labels are enough. The terms (as of September 2026) forbid using outputs to train competing models. OpenAI makes exceptions for undistributed classifiers and embeddings and for fine-tuning its own models, and Anthropic allows it only with express approval. The terms don’t say whether a narrow model for your own use “competes”, so that is a question for a lawyer. An open teacher whose licence allows distillation, such as DeepSeek-R1 under MIT, avoids the question.
- **Does a bigger teacher always make a better student?** No. In Gemma 3’s ablation the smaller teacher won with short training and the larger one only with long training. Models up to 3B don’t consistently gain from strong teachers’ long reasoning traces, and a mix of short and long traces helps them (Li et al., 2025). Pick the teacher by the student’s score on the eval, not by the teacher’s own score.

### Sources

- [Hinton, Vinyals, Dean: Distilling the Knowledge in a Neural Network (2015)](https://arxiv.org/abs/1503.02531)
- [DeepSeek-AI: DeepSeek-R1 (2025), section on distillation into Qwen and Llama](https://arxiv.org/abs/2501.12948)
- [Qwen Team: Qwen3 Technical Report (2025), strong-to-weak distillation](https://arxiv.org/abs/2505.09388)
- [Thinking Machines: On-Policy Distillation (October 2025)](https://thinkingmachines.ai/blog/on-policy-distillation/)
- [Anthropic: Detecting and preventing distillation attacks (February 2026)](https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks)

Interactive page: https://howaiworks.dev/distillation/
