## Vectors and matrices

*How a model reads and predicts*

*Last edited: 28 September 2026*

Every token becomes a vector of numbers that passes through dozens of layers. A layer is mostly that vector multiplied by large, fixed weight matrices, which is why a model’s cost is counted in parameters, bytes and operations.

**In plain words:** Picture a mixing desk with billions of faders. Training set the faders once, and nobody touches them afterwards. Every word passes through the same desk and comes out slightly changed.

*Interactive widget on the page: Click a number in the new vector to see which multiplications produced it.*

### A token’s path through the model

- The embedding is a plain lookup, not a multiplication: the token ID selects a row of a table. In Llama 3 70B the table is 128,256 × 8192 numbers, over 1 billion parameters.
- A layer has two steps. Attention mixes information between tokens (see “Attention”). The MLP processes each token on its own: it expands the vector from 8192 to 28,672 numbers, passes it through a non-linearity (SwiGLU in Llama) and projects it back down. Without the non-linearity, consecutive multiplications could be collapsed into a single matrix. In Llama 3 70B the MLP holds over 80% of a layer’s parameters.
- Each step adds its output to the vector instead of replacing it: `x = x + Attn(Norm(x))`, then `x = x + MLP(Norm(x))`. This vector is the residual stream, a shared bus that layers read from and write to. The addition gives the gradient a straight path through dozens of layers, and normalisation (RMSNorm in Llama) keeps the scale of the numbers in check.
- At the end, the last token’s vector is multiplied by an 8192 × vocabulary-size matrix. That yields one logit per token, and a softmax turns the logits into probabilities (see “The next token”).

*Interactive widget on the page: Click a concept to see its nearest neighbours in vector space.*

### Where the knowledge lives

- Inside there is no database of facts and no if-then rules, only numbers in matrices. Research suggests that MLP layers act as key–value memory (Geva et al., 2021): a pattern in the input “lights up” a direction that writes the associated information into the vector.
- A fact doesn’t sit in one place. Interpretability research shows that a model packs in more features than it has dimensions by overlapping them, so a single number rarely means one thing. That is why you can’t browse a model’s knowledge or reliably fix a single fact. New or changing knowledge goes into the context (see “RAG”).

### What this means for cost

- Memory for the weights is parameters × bytes per parameter: a 70B model takes about 140 GB in BF16 and just under 40 GB at 4 bits (see “Quantisation”).
- Compute: about 2 operations (a multiply and an add) per parameter per token at inference, and about 6 in training because of the backward pass. A 70B model needs about 140 GFLOP per token, plus attention, whose cost grows with context and in Llama 3 70B catches up with the weights at about 100k tokens (see “Attention”). In MoE only the active parameters count (see “Mixture of Experts”).
- Matrix multiplication is thousands of independent dot products, so it suits GPUs perfectly. Yet when generating a single token, the card mostly waits for the weights to be read from memory (see “Why the GPU is idle”).

### Check yourself

**Question:** What happens to a token inside the model, and what does one token cost?

**Short answer:** The token ID selects a vector from the embedding table. The vector passes through dozens of blocks: in each, attention mixes in information from earlier tokens and an MLP processes the token on its own. Each step’s output is added to the vector (the residual stream), with normalisation before the step. At the end, a multiplication by the vocabulary matrix yields logits, and a softmax gives the next-token distribution. At short context almost all the cost is multiplication by fixed weight matrices: about 2 operations per parameter per token, so a 70B model needs about 140 GFLOP per token and its BF16 weights take about 140 GB. Attention adds a cost that grows with context length and in a 70B model catches up with the weights at about 100k tokens.

### Follow-up questions

- **How much compute and memory does one token cost?** About 2 operations per active parameter at inference and 6 in training, plus attention, which grows with context length. Memory is the weights (parameters × bytes) plus the KV cache, which grows with context length (see “The generation loop and KV cache”).
- **Where does a model’s knowledge physically live?** In the weights, largely in the MLP layers, which act as associative memory. Attention moves information between positions. Facts are spread out and overlap one another, which is why it is cheaper to supply new knowledge in context than to write it into the weights (see “RAG” and “Fine-tuning and LoRA”).
- **Why residual connections and normalisation?** Adding the output to the input gives the gradient a straight path through dozens of layers, and each layer only has to learn a correction. Normalisation before every step (usually RMSNorm today) keeps the scale of the activations in check. Without them, deep models train unstably.
- **How does a token embedding differ from an embedding for search?** A token embedding is a table row, with no context. An embedding model runs a whole passage through a transformer and returns one vector per text, trained so that texts with similar meaning land close together (see “Embeddings and vector search”).

### Sources

- [Geva et al.: Transformer Feed-Forward Layers Are Key-Value Memories (2021)](https://arxiv.org/abs/2012.14913)
- [Anthropic: Toy Models of Superposition (2022)](https://transformer-circuits.pub/2022/toy_model/index.html)
- [Transformer Explainer: a real GPT-2 in the browser](https://poloclub.github.io/transformer-explainer/)
- [3Blue1Brown: neural networks and transformers](https://www.3blue1brown.com/topics/neural-networks)
- [LLM Visualization in 3D (bbycroft)](https://bbycroft.net/llm)

Interactive page: https://howaiworks.dev/vectors-and-matrices/
