## Tokens

*How a model reads and predicts*

*Last edited: 28 September 2026*

A model never sees letters or words, only a sequence of integers. Each integer is the ID of a piece of text from a fixed vocabulary. Price, context limits and speed are all counted in tokens, and the way text gets split explains several classic model mistakes.

**In plain words:** A model reads like someone who knows syllables and whole common words but has never seen individual letters. Familiar words it takes in whole; rare ones it assembles from pieces.

*Interactive widget on the page: Pick a tokeniser and an example, or type your own text. See where the token boundaries fall and how many characters one token covers.*

### How the vocabulary is built

- The vocabulary is built by BPE (byte-pair encoding). It starts from the 256 possible bytes and, over a large corpus, repeatedly merges the most frequent adjacent pair into a new token until it reaches the target size. Common words end up as a single token; rare ones are assembled from pieces.
- Before BPE runs, a regular expression splits the text into words, numbers and punctuation, so merges never cross those boundaries. A space sticks to the word that follows it. In the GPT-4o tokeniser “strawberry” with a leading space is one token, but without the space, e.g. at the start of the text, it is three: st|raw|berry (type the word alone into the widget).
- BPE works on UTF-8 bytes, so any text can be encoded and there is no “unknown word” token. Rare combinations fall apart into bytes: in GPT-4o the Polish word “Źdźbło” with a leading space is six tokens, and the space plus “Ź” splits into two pieces, neither of which is a whole character. SentencePiece-based tokenisers get the same guarantee through byte fallback.
- Vocabulary size is a design decision: GPT-4 has about 100k tokens, GPT-4o about 200k, Llama 3 128k (100k from the OpenAI tokeniser plus 28k for languages other than English), Gemma 3 262k. A larger vocabulary gives shorter sequences but a bigger embedding table and output layer.
- The vocabulary also has special tokens: end of text, role markers, tool calls. They are separate IDs that a well-built API won’t let you produce by typing ordinary text. A conversation is one token sequence with such markers (see “How a model sees a chat”).

### Effects you see in answers

- Letters: in “How many r’s are in strawberry?” the model sees one ID instead of ten letters and has to remember the spelling from training. Counting letters, rhyming, reversing words and anagrams are unreliable as a result. Reasoning models do better because they first spell the word out letter by letter, and each letter becomes its own token.
- Numbers: GPT-4 and GPT-4o split digits into groups of three from the left (1234567 becomes 123|456|7), Gemma 3 into single digits. With groups, the decimal places of two numbers don’t line up, so “mental” arithmetic is unreliable. For calculations, give the model a tool or a code interpreter.
- A prompt that ends with a space breaks the token boundary. In training data a space almost always belongs to the next word, so the model gets a rare pattern and answers worse. Don’t end a raw completion prompt, or a prefilled start of the answer, with a space. This matters mostly when self-hosting: a chat API starts the answer in a new turn, and current Claude models reject prefill altogether (as of September 2026).

### Cost and limits

- Price, the context window, `max_tokens`, per-minute rate limits and generation speed are all counted in tokens. Languages other than English use more of them: across the whole text of this guide, the Polish version has about 1.4 times as many tokens as the English one with the GPT-4o tokeniser (3.2 versus 4.5 characters per token) and about 1.6 times as many with Llama 3 (the widget shows the counts). The same content therefore costs more, fills the window faster and takes longer to generate.
- The tokeniser belongs to the model. The same text is a different number of tokens across providers, and sometimes across versions of one model: according to Anthropic, the new tokeniser in Claude Opus 4.7 (April 2026) turns the same text into up to about 35% more tokens than Opus 4.6, at an unchanged per-token price. Don’t compare per-million-token prices directly. Count tokens on your own data with the model’s tokeniser or the provider’s endpoint, e.g. `count_tokens` in the Claude API.

### Check yourself

**Question:** How does BPE tokenisation work, and how does it affect a model’s cost and behaviour?

**Short answer:** A tokeniser splits text into pieces from a fixed vocabulary and maps them to integers. BPE builds the vocabulary from 256 bytes by repeatedly merging the most frequent adjacent pair until it reaches the target size, usually 100–260k. Because it works on bytes, any text can be encoded. Common words become one token; rare and non-English words split into pieces, so Polish takes about 1.3–1.6 times as many tokens as English, depending on the tokeniser. The model never sees individual letters or digits, hence mistakes in letter counting and arithmetic. A larger vocabulary shortens sequences but grows the embedding table and the output layer.

### Follow-up questions

- **Why not tokenise by character or by whole word?** Characters make sequences several times longer, and the cost of attention and the KV cache grows with length. Whole words need a gigantic vocabulary, and a typo, a new word or a proper name gets no ID. Byte-level BPE combines short sequences with full coverage.
- **What does a larger vocabulary change?** Shorter sequences, so cheaper context and faster generation, especially in languages other than English. The price is a bigger embedding table and output layer (about 1 billion parameters each in Llama 3 70B) and rare tokens that saw few examples in training.
- **Can you swap the tokeniser of a trained model?** In practice, no. The embedding table and the output layer are tied to specific IDs, so a new vocabulary needs at least retraining those layers, and usually continued pretraining. Adding a few special tokens during fine-tuning is possible, but their embeddings have to be trained (see “Fine-tuning and LoRA”).
- **How would you estimate the cost of a new feature before launch?** On a sample of real data: input and tool definitions with the model’s own tokeniser or the provider’s token-counting endpoint, and output from the usage field of a few real calls. Reasoning models bill thinking tokens as output even when you don’t see them, and no tokeniser can count those in advance (see “Reasoning models”). Price a repeated prefix at the cache rate (see “Prompt caching”), then multiply by the number of calls. Per-million-token prices can’t be compared directly across providers, because the same text is a different number of tokens for each of them.

### Sources

- [Sennrich et al.: Neural Machine Translation of Rare Words with Subword Units (BPE, ACL 2016)](https://arxiv.org/abs/1508.07909)
- [Anthropic: Introducing Claude Opus 4.7 (new tokeniser, April 2026)](https://www.anthropic.com/news/claude-opus-4-7)
- [Andrej Karpathy: Let’s build the GPT Tokenizer](https://www.youtube.com/watch?v=zduSFxRajkE)
- [Tiktokenizer: real tokenisers in the browser](https://tiktokenizer.vercel.app/)

Interactive page: https://howaiworks.dev/tokens/
