Home / Chapter 1 · How a model reads and predicts
Last edited · 5 min read
Vectors and matrices
Every token becomes a vector of numbers that passes through dozens of layers. A layer is mostly that vector multiplied by large, fixed weight matrices, which is why a model’s cost is counted in parameters, bytes and operations.
In plain wordsPicture a mixing desk with billions of faders. Training set the faders once, and nobody touches them afterwards. Every word passes through the same desk and comes out slightly changed.
Click a number in the new vector to see which multiplications produced it
A token’s path through the model
- The embedding is a plain lookup, not a multiplication: the token ID selects a row of a table. In Llama 3 70B the table is 128,256 × 8192 numbers, over 1 billion parameters.
- A layer has two steps. Attention mixes information between tokens (see “Attention”). The MLP processes each token on its own: it expands the vector from 8192 to 28,672 numbers, passes it through a non-linearity (SwiGLU in Llama) and projects it back down. Without the non-linearity, consecutive multiplications could be collapsed into a single matrix. In Llama 3 70B the MLP holds over 80% of a layer’s parameters.
- Each step adds its output to the vector instead of replacing it:
x = x + Attn(Norm(x)), thenx = x + MLP(Norm(x)). This vector is the residual stream, a shared bus that layers read from and write to. The addition gives the gradient a straight path through dozens of layers, and normalisation (RMSNorm in Llama) keeps the scale of the numbers in check. - At the end, the last token’s vector is multiplied by an 8192 × vocabulary-size matrix. That yields one logit per token, and a softmax turns the logits into probabilities (see “The next token”).
Click a concept to see its nearest neighbours in vector space
Real vectors have anywhere from a few hundred to over ten thousand dimensions. Directions carry meaning too: in classic word2vec, the vector “king − man + woman” lands closest to “queen” once the input words themselves are excluded (otherwise “king” usually wins).
An illustrative map: the positions are placed by hand, not computed by a model. The embedding row is the same for a river “bank” and a savings “bank”. Meaning from context is added only by later layers. Vectors for search come from a separate model, one per whole passage of text (see “Embeddings and vector search”).
Where the knowledge lives
- Inside there is no database of facts and no if-then rules, only numbers in matrices. Research suggests that MLP layers act as key–value memory (Geva et al., 2021): a pattern in the input “lights up” a direction that writes the associated information into the vector.
- A fact doesn’t sit in one place. Interpretability research shows that a model packs in more features than it has dimensions by overlapping them, so a single number rarely means one thing. That is why you can’t browse a model’s knowledge or reliably fix a single fact. New or changing knowledge goes into the context (see “RAG”).
What this means for cost
- Memory for the weights is parameters × bytes per parameter: a 70B model takes about 140 GB in BF16 and just under 40 GB at 4 bits (see “Quantisation”).
- Compute: about 2 operations (a multiply and an add) per parameter per token at inference, and about 6 in training because of the backward pass. A 70B model needs about 140 GFLOP per token, plus attention, whose cost grows with context and in Llama 3 70B catches up with the weights at about 100k tokens (see “Attention”). In MoE only the active parameters count (see “Mixture of Experts”).
- Matrix multiplication is thousands of independent dot products, so it suits GPUs perfectly. Yet when generating a single token, the card mostly waits for the weights to be read from memory (see “Why the GPU is idle”).
Check yourself
What happens to a token inside the model, and what does one token cost?
The token ID selects a vector from the embedding table. The vector passes through dozens of blocks: in each, attention mixes in information from earlier tokens and an MLP processes the token on its own. Each step’s output is added to the vector (the residual stream), with normalisation before the step. At the end, a multiplication by the vocabulary matrix yields logits, and a softmax gives the next-token distribution. At short context almost all the cost is multiplication by fixed weight matrices: about 2 operations per parameter per token, so a 70B model needs about 140 GFLOP per token and its BF16 weights take about 140 GB. Attention adds a cost that grows with context length and in a 70B model catches up with the weights at about 100k tokens.
Po polsku
Numer tokena wybiera wektor z tabeli embeddingów. Wektor przechodzi przez kilkadziesiąt bloków: w każdym attention miesza informacje z wcześniejszymi tokenami, a MLP przetwarza token osobno. Wynik każdego kroku dodaje się do wektora (residual stream), z normalizacją przed krokiem. Na końcu mnożenie przez macierz słownika daje logity, a softmax rozkład następnego tokena. Przy krótkim kontekście prawie cały koszt to mnożenie przez stałe macierze wag: ok. 2 operacje na parametr na token, więc model 70B to ok. 140 GFLOP na token, a jego wagi w BF16 zajmują ok. 140 GB. Attention dokłada koszt rosnący z długością kontekstu, który w modelu 70B dogania wagi przy ok. 100 tys. tokenów.
Follow-up questions (4)
- How much compute and memory does one token cost?
- About 2 operations per active parameter at inference and 6 in training, plus attention, which grows with context length. Memory is the weights (parameters × bytes) plus the KV cache, which grows with context length (see “The generation loop and KV cache”).
- Where does a model’s knowledge physically live?
- In the weights, largely in the MLP layers, which act as associative memory. Attention moves information between positions. Facts are spread out and overlap one another, which is why it is cheaper to supply new knowledge in context than to write it into the weights (see “RAG” and “Fine-tuning and LoRA”).
- Why residual connections and normalisation?
- Adding the output to the input gives the gradient a straight path through dozens of layers, and each layer only has to learn a correction. Normalisation before every step (usually RMSNorm today) keeps the scale of the activations in check. Without them, deep models train unstably.
- How does a token embedding differ from an embedding for search?
- A token embedding is a table row, with no context. An embedding model runs a whole passage through a transformer and returns one vector per text, trained so that texts with similar meaning land close together (see “Embeddings and vector search”).