Home / Chapter 8 · Choosing a model
Last edited · 5 min read
Open-weight models
Open-weight means you can download the weights, i.e. the model’s trained matrices, and run it on your own hardware. It is not the same as open source: you usually don’t get the data or the training code, and the licence may restrict use.
In plain wordsA closed model is a restaurant: you eat what they serve and can’t look into the kitchen. Open-weight is a takeaway: you can reheat and season it at home, but you don’t know the recipe. Open source is the dish with the recipe.
Switch the model type and see what comes in the package and what follows from it
What you get and what you don’t
- Weights, inference code and the tokeniser are enough to run, quantise and fine-tune the model, for example with LoRA (see “Fine-tuning and LoRA”). They are not enough to reproduce it or check what it was trained on.
- The OSI’s Open Source AI Definition (version 1.0, 2024) requires the weights, the complete training and inference code, and a description of the data detailed enough for a skilled person to build a similar system, all under OSI-approved terms: use for any purpose without asking permission. The data itself need not be published, but the public and third-party training data must be listed with where to get it. Few models meet it, e.g. Ai2’s OLMo; licences with use restrictions, such as Llama’s, fail it on terms alone.
- Licences vary a lot. Apache 2.0 and MIT (e.g. gpt-oss, Qwen3, DeepSeek-R1) allow almost anything. The Llama licences (3.1 and 4) require a separate licence from Meta if your products had more than 700 million monthly active users on the model’s release date, plus a “Built with Llama” notice. For the multimodal models (Llama 3.2 Vision, Llama 4 Scout and Maverick) the acceptable use policy does not grant the licence at all to individuals and companies based in the EU; end users of a product built on them are exempt.
Why companies want it
- Data never leaves your infrastructure: your own cloud, your own data centre, even an air-gapped network. In regulated industries this is sometimes the only acceptable route.
- The version is frozen: nobody swaps the weights, settings or serving stack without you knowing (see “Compute and “getting dumber””). This holds only if you host it yourself.
- Full control over inference: fine-tuning, your own quantisation, access to logits, any kind of output-format enforcement, no rate limits imposed by a provider.
Set the API price, the node cost, the monthly volume and what one node can serve. See at what load your own GPUs start to pay off
Illustrative. Node throughput depends heavily on the model and the input/output mix: 10 billion tokens a month is about 3,900 tokens per second non-stop, while DeepSeek reports about 15k output or 74k input tokens per second per 8-GPU node for its 671B MoE model (February 2025). The API price is an average of input tokens and the more expensive output tokens. The chart leaves out the people needed to run it, usually the largest cost of self-hosting.
Cost and pitfalls
- A node costs the same at 5% and at 90% load, and traffic has peaks and troughs. Real average load is far below 100%, so you calculate the cost at that load, not at full load. On top of that come people: serving (vLLM, SGLang), monitoring, updates, security. Load weights as safetensors from a trusted publisher at a pinned revision: a pickle checkpoint can run code when loaded, and
trust_remote_coderuns the repository’s Python on your servers. - A third route is the same open model through a third-party host’s API: no GPUs of your own, and usually cheaper than the closed frontier models. Data leaves your infrastructure again, and the host may change the quantisation or the serving stack.
- The same model behaves differently at different hosts: different quantisation, a different chat template, different handling of tool calling and long context. Moonshot measured this for Kimi K2 (November 2025): tool-call schema accuracy ranged from 100% on the official API to 72% at one host, because of outdated serving versions, malformed tool-call IDs and no constrained decoding. Test the specific endpoint with your own eval set, not the model name (see “Evals” and “LLMs in production”).
- On the hardest tasks the closed frontier models usually lead the open ones. The gap can be small and shifts with every release, so a measurement on your task decides, not a leaderboard.
Check yourself
Open-weight vs open source: what’s the difference, and when should you self-host?
Open-weight means you can download the weights and run the model yourself. Open source, per the OSI definition, also requires the training code, a detailed description of the data and terms without use restrictions, which is rare; weights licences often restrict use, for example by scale or region. Self-hosting gives control: data stays in-house, the version never changes, and you can fine-tune and quantise. You pay for GPUs and people, and a node costs the same idle or full, so it pays off at high, steady volume or under hard regulatory constraints. The middle path is an open model from a third-party host.
Po polsku
Open-weight znaczy, że wagi można pobrać i uruchomić model u siebie. Open-source według definicji OSI wymaga też kodu treningu, szczegółowego opisu danych i warunków bez ograniczeń użycia, a to rzadkość; licencje wag często ograniczają użycie, np. skalą albo regionem. Self-hosting daje kontrolę: dane zostają w firmie, wersja się nie zmienia, można dotrenować i skwantyzować. Płacisz za GPU i ludzi, a węzeł kosztuje tyle samo pusty i pełny, więc to się opłaca przy dużym, stałym ruchu albo twardych wymaganiach regulacyjnych. Pośrednia droga to otwarty model u zewnętrznego hosta.
Follow-up questions (4)
- What do you check in a licence?
- Commercial use, scale thresholds (e.g. 700 million monthly active users in Llama), geographic restrictions (the multimodal Llama models are not licensed to individuals or companies based in the EU), the list of prohibited uses, terms for derivative models and for training other models on the outputs, and attribution requirements.
- How do you compare a host of an open model with the original?
- With the same eval set against the specific endpoint, focusing on tool calling, format enforcement and long context, because that is where differences in quantisation and chat template show up first. Also record the quantisation and version the host declares.
- When would you advise against self-hosting despite a lower price per token?
- When traffic is irregular and the node sits idle most of the time, when the team has nobody to run the serving stack, or when the task needs the quality of a frontier closed model. Calculate cost at the real average load, not at 100%.
- A client requires that data never leaves the EU. What are your options?
- A closed model through a cloud with an EU region and a data processing agreement, an open model at a host in the EU, or self-hosting. The choice depends on whether a contract and a region are enough or full control over the infrastructure is required.