Home / Chapter 2 · Where a model’s knowledge comes from
Last edited · 10 min read
Fine-tuning and LoRA
Fine-tuning is further training of an existing model on your examples: it is good at changing how the model answers and poor at changing what it knows. LoRA does it cheaply by freezing the model and training a small correction, often under 1% of the weights.
In plain wordsFine-tuning is on-the-job training for an experienced employee. Afterwards they write in the company’s format and tone, but they won’t memorise the price list; that’s what the binder on the desk is for, i.e. RAG. LoRA is corrections on a transparent sheet laid over the textbook: the original stays untouched, and the sheet can be lifted off or swapped for another.
Pick a problem and see which rung to start from
What it changes and what it doesn’t
- Fine-tuning shifts the weights so that answers resemble your examples. It is best at teaching what shows up in every answer: format, style and tone, consistent behaviour (when to ask a clarifying question, when to refuse) and a narrow skill such as classification or extracting fields from a document. It also transfers a large model’s behaviour on one task to a smaller, cheaper and faster one, trained on the large model’s answers (see “Distillation”).
- Facts go into the weights poorly and with risk. Gekhman et al. (2024): examples with knowledge the model didn’t have are learned more slowly than the rest, and once learned, they linearly increase the tendency to hallucinate. Ovadia et al. (2023): RAG beat continued training on raw text, for both known and new knowledge. A fact in the weights also has no source and can’t be corrected without another training run.
- The order: prompt and context (instructions, examples, structured output), then RAG for knowledge (see “RAG”), fine-tuning last. You change a prompt in a minute and update RAG by adding a document, while fine-tuning means data, training, evals and a new model to maintain. Reach for it when a polished prompt still loses on evals, or when a long prompt full of examples is too expensive at your volume.
Types and where to train
- SFT (supervised fine-tuning): input → reference answer pairs. The loss is computed only on the answer tokens, as in the SFT stage described in “Training”, but on your examples. The model imitates the references, so their quality sets the ceiling.
- Preference tuning, e.g. DPO (Rafailov et al., 2023): pairs of a better and a worse answer to the same prompt. The model raises the probability of the better one relative to the worse one, with no separate reward model and no RL loop. It suits tone and things to avoid, because a difference is easier to show than a perfect reference is to write.
- Reinforcement fine-tuning: the model generates its own answers and a grader scores them. The grader is a function in code, a comparison with the expected result, or a judge model. No reference answers are needed, only a way to score, so it suits tasks with a verifiable result. The model will exploit a leaky grader (reward hacking, see “Training”).
- Hosted fine-tuning of closed models: you upload a JSONL file of examples, the provider trains and serves the model, and you don’t get the weights. Google Vertex AI offers SFT, preference tuning and RL for Gemini models (September 2026). OpenAI is winding down its platform: since 7 May 2026 it hasn’t accepted organisations that never fine-tuned, since 2 July 2026 it also blocks those with no inference on a fine-tuned model in the last 60 days, and from 6 January 2027 no one will be able to create a new training job.
- An open-weight model (see “Open-weight models”) can be trained through a managed service: Together AI takes a file of examples, and Thinking Machines’ Tinker gives you an API for your own SFT, DPO or RL loop on LoRA. They run the GPUs, and you download the adapter and serve it anywhere. Training on your own hardware gives full control of the pipeline, but the GPUs and serving are on you.
How LoRA works
- Full fine-tuning updates every weight. LoRA (Hu et al., 2021) freezes a d × k matrix W and learns only a correction ΔW = B·A, where B is d × r and A is r × k. The rank r (e.g. 8, 16, 64) is much smaller than d and k, so instead of d·k parameters you train r·(d + k): for a 4096 × 4096 matrix and r = 16 that is 131k instead of 16.8 million. The authors assume that the weight change needed to adapt a model has low rank. B starts at zero, so at first ΔW = 0 and the model behaves like the base, and the output is scaled by α/r.
- The original paper added LoRA mainly to attention. Schulman (Thinking Machines, 2025) shows that LoRA on all layers, especially the MLP, matches full fine-tuning on small and medium datasets, and in RL even at rank 1. It loses when there is more data than the adapter can hold, and it tolerates large batches worse than full fine-tuning. The optimal learning rate is about 10 times higher than for full fine-tuning.
- QLoRA (Dettmers et al., 2023) is LoRA on a base quantised to 4 bits in the NF4 format (see “Quantisation”). Gradients flow through the frozen base into a 16-bit adapter. The base takes about 4 times less memory, and in the paper fine-tuning a 65B model fitted on a single 48 GB card with no loss of quality compared with 16-bit training.
Change the rank and the model class. See how many weights LoRA trains and how much GPU memory it needs
GPU memory for weights and optimiser state:
Illustrative numbers. Architecture as in Llama 3.1 8B and 70B, LoRA on all seven linear matrices in every layer (Q, K, V, O and three in the MLP). Memory by rule of thumb: full fine-tuning with Adam in mixed precision is about 16 bytes per parameter (weights and gradients in BF16, an FP32 copy of the weights, two Adam moments), LoRA is 2 bytes per frozen weight in BF16, QLoRA about 0.5 bytes (4 bits), plus 16 bytes for every adapter parameter. Activations come on top: they grow with sequence length and batch size. Cards chosen with 20% headroom for activations.
- An adapter is a small file: for an 8B-class model with r = 16, about 42 million parameters, i.e. about 84 MB in BF16 against a 16 GB base. You keep one base and many adapters, e.g. per customer, language or task. vLLM keeps one copy of the base and mixes requests for different adapters in one batch; S-LoRA (Sheng et al., 2023) serves thousands of adapters on a single GPU this way. The price is an extra B·(A·x) multiplication at every step.
- Merging: after training, B·A can be added to W. The model is then exactly as fast as the base, but every variant is a separate full copy. Several adapters on the same base can also be combined into one without training, by a weighted sum or with TIES and DARE (
add_weighted_adapterin Hugging Face PEFT). Skills can interfere with each other in the process, so evaluate a merged adapter like a new model.
Data, evaluation, maintenance
- Data quality beats quantity. In LIMA (Zhou et al., 2023), 1000 carefully chosen SFT examples were enough for a 65B model to give high-quality answers. The authors conclude that knowledge comes from pretraining and tuning mostly teaches form. The model learns everything that repeats in the data, including typos, inconsistent formatting and bad answers.
- Set the eval set aside before training and never train on it (see “Evals”). First measure the base with your best prompt on it: that is the bar fine-tuning has to clear. Add out-of-task cases and refusal tests, because fine-tuning weakens safeguards: Qi et al. (2023) stripped them from GPT-3.5 Turbo with 10 examples for less than $0.20, and ordinary, benign datasets weakened them too, just less.
- Overfitting: training loss falls, validation loss rises, and the model parrots phrases from the examples. Fewer epochs, a lower learning rate and picking the checkpoint by validation loss help. Catastrophic forgetting: the model loses skills outside your data. LoRA forgets less than full fine-tuning, but on large datasets it also learns less (Biderman et al., 2024, on code and maths: the capacity-limited case above).
- A fine-tuned model is tied to its base. An adapter fits only that exact version of the weights, and with a hosted provider the model disappears along with its base: on 23 October 2026 OpenAI shuts down ft-gpt-4.1-nano, among others. A new base means a new training run, so you version the data, configuration, base version, adapter and eval results together, and run the whole pipeline with one command. With every new base, first check whether a prompt alone is now enough.
Check yourself
When would you use fine-tuning instead of a prompt or RAG, and how does LoRA work?
Order: prompt and context first, RAG for knowledge, fine-tuning last. Fine-tuning is good at behaviour: format, tone, a narrow task, or distilling a large model’s behaviour into a small one. It is poor at facts: they go in unreliably, without a source, every change means retraining, and they can increase confident errors. LoRA freezes the weights and learns a low-rank update ΔW = B·A of rank r, so r·(d + k) parameters instead of d·k. An adapter is tens of MB, so many adapters can be served on one base. QLoRA does the same on a 4-bit base.
Po polsku
Kolejność: prompt i kontekst, dla wiedzy RAG, fine-tuning na końcu. Fine-tuning dobrze uczy zachowania: formatu, tonu, wąskiego zadania, albo przenosi zachowanie dużego modelu do małego. Faktów uczy słabo: wchodzą niepewnie, bez źródła, każda zmiana to nowy trening, a do tego mogą zwiększyć liczbę pewnych siebie błędów. LoRA zamraża wagi i uczy poprawki ΔW = B·A o niskim ranku r, czyli r·(d + k) parametrów zamiast d·k. Adapter waży dziesiątki MB, więc na jednej bazie serwuje się wiele adapterów. QLoRA robi to samo na bazie w 4 bitach.
Follow-up questions (5)
- You have 50k company documents that change every week. Fine-tuning or RAG?
- RAG. Facts from fine-tuning go in unreliably, have no source, and every change needs a new training run. Fine-tuning can come later to teach the model the answer format and how to cite passages, but the knowledge stays in the index.
- How do you choose the LoRA rank and the layers to apply it to?
- All linear matrices, including the MLP, because attention-only LoRA clearly lags behind. Rank is the adapter’s capacity: on small and medium datasets a low rank is enough, in RL even r = 1, while on large datasets LoRA starts losing to full fine-tuning. The learning rate is about 10 times higher than for full fine-tuning; compare a few ranks on the eval.
- You have 200 customers and each wants a model in their own style. How do you serve that?
- One base and one LoRA adapter per customer, without merging. vLLM picks the adapter for each request, keeps several active adapters in one batch next to a single copy of the base and loads more on the fly. The price is a little extra compute at every step instead of 200 full copies of the model.
- After fine-tuning the format is perfect, but the model does worse outside the task and is easier to talk into things it used to refuse. What happened?
- Catastrophic forgetting and weakened safeguards: training pushed the weights only towards your examples. Fewer epochs, a lower learning rate or LoRA instead of full fine-tuning help, as do general examples and refusals in the data, with an out-of-task set and safety tests in the eval.
- The provider retires the base model your fine-tune is built on. What do you do?
- First check whether the new base with just a prompt already passes the eval, because then fine-tuning is no longer needed. If not, retrain on the new base with the same pipeline. Data, configuration and the eval set are versioned, so it is a single run.
Sources
- Hu et al.: LoRA: Low-Rank Adaptation of Large Language Models (2021)
- Dettmers et al.: QLoRA: Efficient Finetuning of Quantized LLMs (2023)
- John Schulman, Thinking Machines: LoRA Without Regret (2025)
- Gekhman et al.: Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations? (2024)
- Thinking Machines: Tinker, an API for training open-weight models with LoRA