An interactive guide for engineers
How AI works, from the inside
How language models work under the hood and what matters when you build agentic systems on top of them. The essence first, then the details, with a widget to play with on every page.
Start from the beginning →1.1 Tokens · 34 topics in 9 chapters · about 4 h
What you will find
Explainers stop at the next token, and agent courses skip what is inside the model. This guide connects the two: why the causal mask lets prompt caching work only on an identical prefix, and why that makes every agent turn cost more.
- Inside the modelTokens, attention, sampling, training and reasoning, explained through the mechanism, not the hype.Chapter 1 · How a model reads and predicts →
- What it costs and whyThe KV cache, context windows, prompt caching and why the GPU spends most of its time waiting.Chapter 3 · How a model writes and what it costs →
- Agentic systemsThe agent loop, tools, MCP, context engineering, RAG and multi-agent set-ups, with their failure modes.Chapter 5 · Agents →
- Quality and productionEvals, security and reliability, plus a short answer and follow-up questions on every page to check yourself.Chapter 7 · Production and serving →
34 topics · English and Polish · every page also as plain text for AI assistants: llms-full.txt