A free, hands-on textbook on LLMs
Understand how language models do the math. Then build your own.
From a single neuron to a Transformer that writes text. No machine learning background needed: every formula comes after a worked example with real numbers, and every chapter has an experiment you can run in your browser.
What you’ll be able to do
- Explain it: What happens inside a Transformer, step by step – in your own words, with small numbers.
- Work it out: Neurons, softmax, attention and backpropagation by hand, until the formulas feel obvious.
- Build it: A complete Python project: tokenizer, model, training and text generation, running on your own machine.
The path in six parts
I · Understanding neural networks
- Your path to your own language model
- What does it mean for a program to learn?
- An artificial neuron you can compute by hand
- Why neurons need an activation function
- From neurons to a small network
- Mathematical symbols in plain language
- How do we measure a wrong answer?
- How numbers actually learn
- Learning through several layers
II · From numbers to language
- Your Python workshop
- Vectors, matrices and tensor shapes
- From text to token IDs
- Language as a probability problem
- Your first real language model
- Embeddings and shared representations
- The learning goal: cross-entropy
III · Building the transformer
- Self-attention without the mystery
- Causality and multiple heads
- The complete transformer block
- Implementing your own decoder
IV · Training and evaluating
- Data is part of the model
- Your first training loop
- Reading the loss and finding bugs
- From scores to text
V · Modern language models
- RMSNorm, RoPE and SwiGLU
- GQA, FlashAttention and the KV cache
- Precision and training memory
- From a text model to an answer model
- Evaluating answers honestly
V · Modern language models
- Scaling up without calculating blindly
- What today's large systems need on top
- Your final project: from zero to an answer
- Glossary, formulas and sources
Frequently asked questions
Do I need a maths background?
Adding and multiplying is enough to get started. Every new piece of notation is explained before you need it, and always worked through with numbers first.
Do I need a GPU?
No. All the experiments run in your browser, and the Python project trains small models on an ordinary laptop. A GPU just makes bigger runs faster.
Will I be able to build ChatGPT afterwards?
You’ll understand the same building blocks and train a small model yourself. Large systems differ mainly in the amount of data, compute and polish – the last part of the book explains how.
Does it cost anything?
No. No account, no cookies, no tracking. Your reading progress stays in your browser.