Home / Chapter 1 · How a model reads and predicts
    Last edited · 5 min read

    Use with AI

    Vectors and matrices

    Every token becomes a vector of numbers that passes through dozens of layers. A layer is mostly that vector multiplied by large, fixed weight matrices, which is why a model’s cost is counted in parameters, bytes and operations.

    In plain wordsPicture a mixing desk with billions of faders. Training set the faders once, and nobody touches them afterwards. Every word passes through the same desk and comes out slightly changed.

    Click a number in the new vector to see which multiplications produced it

    Scale: in Llama 3 70B a token vector has 8192 numbers, and the largest matrix in a layer is 8192 × 28,672, i.e. 235 million multiplications. There are 80 layers of 7 matrices each, so every token costs about 70 billion multiply-adds.

    A token’s path through the model

    Click a concept to see its nearest neighbours in vector space

    Real vectors have anywhere from a few hundred to over ten thousand dimensions. Directions carry meaning too: in classic word2vec, the vector “king − man + woman” lands closest to “queen” once the input words themselves are excluded (otherwise “king” usually wins).

    An illustrative map: the positions are placed by hand, not computed by a model. The embedding row is the same for a river “bank” and a savings “bank”. Meaning from context is added only by later layers. Vectors for search come from a separate model, one per whole passage of text (see “Embeddings and vector search”).

    Where the knowledge lives

    What this means for cost

    Check yourself

    What happens to a token inside the model, and what does one token cost?

    The token ID selects a vector from the embedding table. The vector passes through dozens of blocks: in each, attention mixes in information from earlier tokens and an MLP processes the token on its own. Each step’s output is added to the vector (the residual stream), with normalisation before the step. At the end, a multiplication by the vocabulary matrix yields logits, and a softmax gives the next-token distribution. At short context almost all the cost is multiplication by fixed weight matrices: about 2 operations per parameter per token, so a 70B model needs about 140 GFLOP per token and its BF16 weights take about 140 GB. Attention adds a cost that grows with context length and in a 70B model catches up with the weights at about 100k tokens.

    Po polsku

    Numer tokena wybiera wektor z tabeli embeddingów. Wektor przechodzi przez kilkadziesiąt bloków: w każdym attention miesza informacje z wcześniejszymi tokenami, a MLP przetwarza token osobno. Wynik każdego kroku dodaje się do wektora (residual stream), z normalizacją przed krokiem. Na końcu mnożenie przez macierz słownika daje logity, a softmax rozkład następnego tokena. Przy krótkim kontekście prawie cały koszt to mnożenie przez stałe macierze wag: ok. 2 operacje na parametr na token, więc model 70B to ok. 140 GFLOP na token, a jego wagi w BF16 zajmują ok. 140 GB. Attention dokłada koszt rosnący z długością kontekstu, który w modelu 70B dogania wagi przy ok. 100 tys. tokenów.

    Follow-up questions (4)
    How much compute and memory does one token cost?
    About 2 operations per active parameter at inference and 6 in training, plus attention, which grows with context length. Memory is the weights (parameters × bytes) plus the KV cache, which grows with context length (see “The generation loop and KV cache”).
    Where does a model’s knowledge physically live?
    In the weights, largely in the MLP layers, which act as associative memory. Attention moves information between positions. Facts are spread out and overlap one another, which is why it is cheaper to supply new knowledge in context than to write it into the weights (see “RAG” and “Fine-tuning and LoRA”).
    Why residual connections and normalisation?
    Adding the output to the input gives the gradient a straight path through dozens of layers, and each layer only has to learn a correction. Normalisation before every step (usually RMSNorm today) keeps the scale of the activations in check. Without them, deep models train unstably.
    How does a token embedding differ from an embedding for search?
    A token embedding is a table row, with no context. An embedding model runs a whole passage through a transformer and returns one vector per text, trained so that texts with similar meaning land close together (see “Embeddings and vector search”).

    Sources

    Report an error · Suggest a fix