Tokenwerk · The LLM textbook

A free, hands-on textbook on LLMs

Understand how language models do the math. Then build your own.

From a single neuron to a Transformer that writes text. No machine learning background needed: every formula comes after a worked example with real numbers, and every chapter has an experiment you can run in your browser.

Start with chapter 1

What you’ll be able to do

  • Explain it: What happens inside a Transformer, step by step – in your own words, with small numbers.
  • Work it out: Neurons, softmax, attention and backpropagation by hand, until the formulas feel obvious.
  • Build it: A complete Python project: tokenizer, model, training and text generation, running on your own machine.

The path in six parts

I · Understanding neural networks

  1. Your path to your own language model
  2. What does it mean for a program to learn?
  3. An artificial neuron you can compute by hand
  4. Why neurons need an activation function
  5. From neurons to a small network
  6. Mathematical symbols in plain language
  7. How do we measure a wrong answer?
  8. How numbers actually learn
  9. Learning through several layers

II · From numbers to language

  1. Your Python workshop
  2. Vectors, matrices and tensor shapes
  3. From text to token IDs
  4. Language as a probability problem
  5. Your first real language model
  6. Embeddings and shared representations
  7. The learning goal: cross-entropy

III · Building the transformer

  1. Self-attention without the mystery
  2. Causality and multiple heads
  3. The complete transformer block
  4. Implementing your own decoder

IV · Training and evaluating

  1. Data is part of the model
  2. Your first training loop
  3. Reading the loss and finding bugs
  4. From scores to text

V · Modern language models

  1. RMSNorm, RoPE and SwiGLU
  2. GQA, FlashAttention and the KV cache
  3. Precision and training memory
  4. From a text model to an answer model
  5. Evaluating answers honestly

V · Modern language models

  1. Scaling up without calculating blindly
  2. What today's large systems need on top
  3. Your final project: from zero to an answer
  4. Glossary, formulas and sources

Frequently asked questions

Do I need a maths background?

Adding and multiplying is enough to get started. Every new piece of notation is explained before you need it, and always worked through with numbers first.

Do I need a GPU?

No. All the experiments run in your browser, and the Python project trains small models on an ordinary laptop. A GPU just makes bigger runs faster.

Will I be able to build ChatGPT afterwards?

You’ll understand the same building blocks and train a small model yourself. Large systems differ mainly in the amount of data, compute and polish – the last part of the book explains how.

Does it cost anything?

No. No account, no cookies, no tracking. Your reading progress stays in your browser.