Home / Chapter 1 · How a model reads and predicts
    Last edited · 6 min read

    Use with AI

    The next token

    A model doesn’t return text, only a score (logit) for every token in the vocabulary. Separate code, the sampler, turns the scores into probabilities and draws one token. Every token of an answer is produced this way, one after another.

    In plain wordsYour phone keyboard suggests three words. A model does the same for a hundred thousand pieces at once, then rolls a die weighted by those odds. Temperature decides how heavily the die is loaded.

    Change the temperature and top-p, then sample. Compare a question the model knows, one it gets confidently wrong and one where it is guessing

    Real next-token probabilities from Qwen3-0.6B-Base, a small open model without chat training: its eight most likely tokens, rescaled to 100%. Struck-through bars have been cut by top-p.

    From logits to a token

    Pitfalls

    What you set in practice

    Check yourself

    How does a model choose the next token, and how do you set temperature and top-p?

    The model returns logits for the whole vocabulary. A temperature softmax, exp(logit/T), turns them into a distribution: T below 1 sharpens it, above 1 flattens it, and T = 0 always picks the top token. Top-p keeps the smallest set of tokens covering, say, 90% of the probability and cuts the tail where odd tokens come from. On non-reasoning models, use a low temperature for extraction and code and a higher one for creative text, change one parameter at a time and check on evals. T = 0 is not fully reproducible, and reasoning models often lock the sampler or, like Gemini 3, work best at the default 1.0.

    Po polsku

    Model zwraca logity dla całego słownika. Softmax z temperaturą, exp(logit/T), zamienia je w rozkład: T poniżej 1 wyostrza, powyżej 1 spłaszcza, a T = 0 to zawsze najlepszy token. Top-p zostawia najmniejszy zbiór tokenów pokrywający np. 90% prawdopodobieństwa i odcina ogon, z którego biorą się dziwne tokeny. W modelach bez rozumowania ekstrakcja i kod dostają niską temperaturę, teksty twórcze wyższą; zmieniaj jeden parametr naraz i sprawdzaj na ewaluacjach. T = 0 nie daje pełnej powtarzalności, a modele rozumujące często blokują sampler albo, jak Gemini 3, działają najlepiej przy domyślnym 1,0.

    Follow-up questions (4)
    Why doesn’t temperature 0 give full reproducibility?
    The result depends on which other requests landed in the same batch: a different batch size means a different order of floating-point operations and slightly different logits. With two nearly tied tokens that is enough to pick the other one, and from there the whole answer diverges.
    Top-k, top-p, min-p: what’s the difference?
    Top-k keeps a fixed number of candidates regardless of how confident the model is. Top-p keeps as many as it takes to cover a given share of the probability, so it adapts to the distribution. Min-p drops tokens weaker than a fraction of the best one, e.g. 0.1 × its probability. Its authors report that it copes better with high temperatures, but a 2025 reanalysis found their evidence doesn’t support that.
    Does a lower temperature reduce hallucinations?
    Only the ones that come from drawing a token from the tail. If the model “believes” a wrong fact, T = 0 will pick it every time (see “Hallucinations”).
    How do you use logprobs for classification?
    Have the model answer with a single label token and read the probability of each label. You get a distribution instead of a single answer, and you set a threshold below which the case goes to a human. Calibrate the threshold on your own data, because after post-training the model’s confidence is often inflated.

    Sources

    Report an error · Suggest a fix