# How AI works > An interactive guide for engineers: how language models work under the hood and what matters when building agentic systems. From intuition to nuance, with questions to check yourself. The guide is in English (Polish version under /pl/). Every topic: intuition, an interactive widget, mechanism, nuances, a short answer to check yourself, follow-up questions and sources. Last updated: 2026-09-26. What’s new: https://howaiworks.dev/changelog/ The whole guide in one file: https://howaiworks.dev/llms-full.txt License: text and images CC BY 4.0 (quote and adapt with a link to the source), code MIT. ## How a model reads and predicts From characters on screen to deciding what the next piece of text will be. - [Tokens](https://howaiworks.dev/tokens.md): A model sees no letters, only pieces of words turned into numbers. - [Vectors and matrices](https://howaiworks.dev/vectors-and-matrices.md): A token is a vector, and a layer multiplies it by fixed weight matrices. - [Attention](https://howaiworks.dev/attention.md): Every word looks back and chooses which earlier words to take information from. - [The next token](https://howaiworks.dev/sampling.md): The model gives odds for every token, and the sampler draws one. ## Where a model’s knowledge comes from Training, distillation, fine-tuning, the assistant role in a chat, and thinking before answering. - [Training](https://howaiworks.dev/training.md): Wide reading, vocational training, graded practice. - [Distillation](https://howaiworks.dev/distillation.md): How a small model learns from a large one: hard labels and distributions, on-policy, reasoning, limits and terms of service. - [Fine-tuning and LoRA](https://howaiworks.dev/fine-tuning.md): When to fine-tune and when a prompt or RAG is enough, how LoRA trains a fraction of the weights, and what fine-tuning breaks. - [How a model sees a chat](https://howaiworks.dev/chat-template.md): A conversation is one document, and the model writes the next line. - [Reasoning models](https://howaiworks.dev/reasoning-models.md): Scratch paper before the answer: it helps on multi-step tasks, and you pay for every line. ## How a model writes and what it costs Writing one piece at a time, caching, and why long conversations are expensive. - [The generation loop and KV cache](https://howaiworks.dev/kv-cache.md): The generation loop, prefill versus decode, and the GPU memory taken up by the key-value cache. - [Prompt caching](https://howaiworks.dev/prompt-caching.md): The KV cache kept by the provider between requests: cheaper input and a faster start, but only for an identical prompt prefix. - [The context window and agents](https://howaiworks.dev/context-window.md): What counts towards the limit, why each agent turn costs more than the last, and why quality drops with length. ## Context and knowledge What to put into the context window and how to find it in your data. - [Context engineering and memory](https://howaiworks.dev/context-engineering.md): What to put in the window before each call, and how an agent remembers what doesn’t fit. - [Embeddings and vector search](https://howaiworks.dev/embeddings.md): Searching by meaning instead of by words: how it works, what it costs and where it fails. - [RAG](https://howaiworks.dev/rag.md): The pipeline from indexing and chunking, through retrieval and reranking, to citations, permissions and evaluation. ## Agents A model in a loop: tools, the standard for connecting them, enforced formats, the coding-agent harness and splitting the work. - [The agent loop](https://howaiworks.dev/agent-loop.md): A model in a loop picks the next step itself, until it decides it is done. - [Tools (function calling)](https://howaiworks.dev/tool-calling.md): The model doesn’t press buttons. It writes which one to press, and your code presses it. - [MCP](https://howaiworks.dev/mcp.md): One standard for connecting tools to many model-powered applications. The model still sees only definitions in the prompt. - [Enforcing output format](https://howaiworks.dev/structured-output.md): Constrained decoding: strict mode guarantees JSON that matches the schema, but not its content. - [The coding-agent harness](https://howaiworks.dev/coding-agent-harness.md): Everything around the model that turns it into a coding agent: tools, instructions, context, permissions and undo. - [Multiple agents](https://howaiworks.dev/multi-agent.md): A lead agent delegates parts of a task to subagents with clean contexts. Faster when the parts are independent, usually more expensive. ## Quality and security Hallucinations, measuring quality instead of guessing, and attacks through content. - [Hallucinations](https://howaiworks.dev/hallucinations.md): A model doesn’t say “I don’t know”; it guesses in the same confident tone. - [Evals](https://howaiworks.dev/evals.md): A fixed set of cases after every change, plus measurement in production, instead of “it seems better”. - [Prompt injection](https://howaiworks.dev/prompt-injection.md): Text the agent only reads can start steering it. Architecture defends against this, not the prompt. ## Production and serving A reliable system around the model, and what happens in the data centre. - [LLMs in production](https://howaiworks.dev/llms-in-production.md): What to build around a model API: retries, fallbacks, limits, costs, traces and safe changes. - [Why the GPU is idle](https://howaiworks.dev/roofline.md): During generation the GPU mostly waits for weights to arrive from memory. That explains batching, quantisation and pricier output tokens. - [Continuous batching](https://howaiworks.dev/continuous-batching.md): The server swaps conversations in and out of the batch at every step and gives them memory in blocks, so batch slots don’t sit empty. - [Speculative decoding](https://howaiworks.dev/speculative-decoding.md): A cheap draft guesses several tokens and the large model checks them in one pass. Faster, with the same output. - [Quantisation](https://howaiworks.dev/quantization.md): Weights in 8 or 4 bits instead of 16: 2–4 times less memory and faster generation for a small loss in quality. - [Mixture of Experts](https://howaiworks.dev/mixture-of-experts.md): Many experts, and a router picks a few per token: compute like a small model, memory like a large one. ## Choosing a model Matching a model to the task, open weights, alternatives to an LLM, and why a model sometimes seems “dumber”. - [Choosing a model](https://howaiworks.dev/choosing-a-model.md): The cheapest model that passes your eval set: cost per task, latency, limits and reading benchmarks. - [Open-weight models](https://howaiworks.dev/open-weight-models.md): Weights you can download and run yourself. Control over data and version, at the cost of GPUs and operations. - [When not to use an LLM: classifiers and System One](https://howaiworks.dev/when-not-to-use-an-llm.md): When a classifier, logprob scoring or a decision model beats an LLM. TypeSafe’s System One as the worked example, with vendor claims kept apart from facts. - [Compute and “getting dumber”](https://howaiworks.dev/compute.md): Pinned weights do not change; the system around them does, and apps and aliases can switch to new weights. What has been documented, what is a hypothesis, and how to measure it. ## Practice and review Design scenarios and flashcards for review. - [Design a system](https://howaiworks.dev/system-design.md): Five design scenarios from real work: requirements, architecture, decisions, evaluation, failures. - [Flashcards](https://howaiworks.dev/flashcards.md): One question for every topic: answer it yourself first, then flip the card.