## Mixture of Experts

*Production and serving*

*Last edited: 28 September 2026*

In a Mixture of Experts (MoE) model, the MLP layer is replaced by many “experts”, and a router picks a few of them for each token. The model has a huge number of parameters but computes with only some of them per token: compute like a small model, memory like a large one.

**In plain words:** A clinic. Reception sends you to two of its eight doctors. They all have to be in the building, but you pay for only two appointments. When a crowd arrives every doctor has patients, and reception has to make sure the whole queue doesn’t end up outside one door.

*Interactive widget on the page: Click the tokens and watch which experts the router picks. At the bottom: how many experts the whole sentence needs.*

### How routing works

- The router is a small linear layer. From the token’s vector it computes a score for each expert, picks the top k (usually 2–8), and sums their outputs with weights derived from those scores. The choice is made separately in every MoE layer and for every token.
- Experts replace only the MLP. Attention, embeddings and normalisation are shared, which is why Mixtral 8×7B has about 47 billion parameters rather than 56 billion, of which about 13 billion are active per token. Newer models have many more, smaller experts: DeepSeek-V3 has 256 routed experts and one shared expert in each MoE layer, and picks 8.
- An “expert” is not a topic specialist. The Mixtral authors found no clear pattern of assignment by the text’s domain; the choice relates more to syntax, e.g. indentation in code consistently goes to the same experts.
- Without load balancing the router collapses onto a few favourite experts, and the rest of the parameters barely learn. Hence an auxiliary balancing loss in training (since the first sparse MoE layers, simplified in Switch Transformer) or, in DeepSeek-V3, mainly a per-expert bias adjusted during training and used only to pick experts, plus a tiny per-sequence balancing loss.

### Consequences for serving

- Every expert has to sit in memory. DeepSeek-V3 has 671 billion parameters, 37 billion of them active per token. In FP8 that is about 671 GB of weights alone, more than the 640 GB on eight H100s.
- With one conversation, decode reads only the active experts, so it is as fast as a model of that active size. With a larger batch, tokens spread to almost every expert: a step reads almost the whole model, and each expert gets only part of the batch. For compute to catch up with reading, the batch has to be larger than for a dense model by roughly the ratio of total experts to selected ones (see “Why the GPU is idle”).
- That is why a large MoE is spread across many GPUs by expert (expert parallelism), with an all-to-all exchange of tokens between GPUs in every layer. DeepSeek-V3 decodes on 320-GPU units, one expert per GPU, and duplicates the most frequently picked experts, because the most loaded expert sets the time of the whole step.

### MoE or a dense model

- MoE gives more quality per unit of compute and cheaper training. A dense model of the same total size is usually better, but computes many times more per token. Most leading open-weight models since 2025 are MoE, which is why two numbers are quoted: total and active parameters (e.g. Qwen3-235B-A22B).
- On hardware with large but slow memory, like a Mac with unified memory, MoE runs surprisingly fast for a single user, because it reads only the active parameters. gpt-oss-120b (117 billion parameters, 5.1 billion active) fits on a single 80 GB GPU thanks to MXFP4. For the same reason a large MoE runs on one consumer GPU if the expert weights stay in CPU RAM and attention runs on the GPU (llama.cpp `--cpu-moe` or `--n-cpu-moe`).
- Self-hosting a large MoE pays off only with enough traffic to fill a batch across many GPUs. With less traffic, an API or a smaller model is cheaper (see “Open-weight models”).

### Check yourself

**Question:** What is MoE, and what are its consequences for serving?

**Short answer:** In MoE the MLP layer is replaced by many experts, and a router picks a few for every token in every layer. Only the active parameters are computed, for example 37 of 671 billion in DeepSeek-V3, so the quality is closer to a large model at the compute of a small one. The price is memory, because every expert must be loaded, and serving: at larger batch sizes tokens hit almost every expert, so it takes much bigger batches, expert parallelism with all-to-all communication between GPUs, and care to keep the expert load balanced.

### Follow-up questions

- **An MoE model has 5 billion active parameters. Will it behave like a 5B model?** It computes like 5B, but the whole model has to sit in memory. With one conversation, decode is as fast as a small model. At a medium batch size each step reads almost all the experts, so it costs as much as reading a large model while doing little computation. Only a very large batch brings the cost per token close to a 5B model.
- **Why balance expert load during training?** Without it the router collapses onto a few favourite experts, which get better and are picked more and more often, while the rest of the parameters go to waste. In serving, uneven load means the most loaded expert sets the step time.
- **How would you spread a large MoE across GPUs?** Attention usually with tensor or data parallelism, experts with expert parallelism: each GPU holds some of the experts, and tokens are exchanged all-to-all in every layer. What matters most is a fast interconnect between GPUs and duplicating the most frequently picked experts. DeepSeek-V3 does this on 320 GPUs for decode.

### Sources

- [Hugging Face: Mixture of Experts Explained](https://huggingface.co/blog/moe)
- [Jiang et al.: Mixtral of Experts (2024)](https://arxiv.org/abs/2401.04088)
- [DeepSeek-AI: DeepSeek-V3 Technical Report (2024)](https://arxiv.org/abs/2412.19437)
- [Maarten Grootendorst: A Visual Guide to Mixture of Experts](https://newsletter.maartengrootendorst.com/p/a-visual-guide-to-mixture-of-experts)

Interactive page: https://howaiworks.dev/mixture-of-experts/
