Tokenwerk · The LLM textbook

Chapter 1 · I · Understanding neural networks · 5 minutes

Your path to your own language model

We start with a single adjustable number. Step by step, it grows into a neural network that learns from text examples.

You don't need to know neural networks yet

This book starts before the technical terms. To begin, all you need is addition, multiplication and a willingness to work through small examples yourself. Whenever mathematical notation comes in, we explain how to read it. Programming experience helps in the later hands-on part. We don't assume any experience with machine learning.

Our goal is a language model of your own: a program that learns from example texts which pieces of text can follow one another. Give it “The cat sits on the”, and it computes possible continuations. From these it can assemble a new text step by step. A small model can learn simple language patterns. Reliable answers to new questions are a further goal, which we train and test separately.

Before we deal with language, we settle a much smaller question: how can a program learn from examples to multiply a number by three, without us telling it the three? This manageable case teaches you the core of training.

What you build step by step

At first our model consists of nothing but one adjustable number. Then we build a computing unit, an artificial neuron. Several such units are connected into a small neural network. We compute an output by hand and look at how wrong outputs can lead to better settings.

Only then do we carry the idea over to language. Text is translated into numeric IDs. The model processes these IDs and scores possible continuations. Later it learns to connect earlier parts of the text in a targeted way. The network architecture used for this is called the transformer. Architecture here simply means the blueprint of the network: which parts it contains and how they are connected.

By the end you will have a complete workflow: prepare texts, create a model with random initial values, train it, save it and use it to generate text. In the final hands-on part you also train on questions and desired answers.

How each chapter is meant to work

Start with the question the chapter wants to answer. Then read the small example and work it through yourself. Only once the numbers make sense, read the same calculation in its shorter mathematical notation. So the formula describes something you have already seen.

Under “Experiment” you can change numbers. Before you move a slider, predict what should happen. Under “Practice” you'll find an open exercise and a comprehension check. Compare your solution only after your own attempt. A wrong answer shows you which spot you should look at again.

At the start of each chapter, the required prior knowledge is listed with links. Terms with a dotted underline can be tapped; a short explanation then opens with an example and, where relevant, a link to the chapter that covers the basics. These explanations supplement the text. You don't have to keep switching to the glossary.

Three things with similar names that are different

Using a model: You give an already trained program an input and get an output. Normally it is not trained any further in the process.

Training an existing model further: You start with the settings it has already learned and adjust them with additional examples. This is called fine-tuning. We introduce the term more precisely later.

Training a model from scratch: You start with random settings. No finished model weights are loaded. This is exactly the path we take. “Weights” are the adjustable numbers in the network; you'll see their effect as early as the neuron chapter.

What “modern” means here

The book takes you as far as building blocks that are common in modern language models. You don't need to learn their names yet. For each new building block, you will later understand which problem it is meant to solve and how it changes the calculation so far.

A small model of your own stays much smaller and more limited than a very large assistant system. Our first goal is therefore one you can check: the model should learn simple patterns and be able to produce understandable short continuations. After that we test whether it can also answer new questions sensibly within a limited subject area. That requires enough suitable data and training time.

What you actually do yourself

The website contains small computational experiments that run directly in your browser. You train the full model later with the downloadable Python project on your own computer. Programming only starts after you have got to know the basic calculations.

Reading progress, answers and notes are stored in this browser. The “Chapter completed” button is your own assessment. If you switch devices, this local progress is not carried over automatically.

A small notebook helps: before each attempt, write down what you want to change and what result you expect. Afterwards, note what actually happened. That alone makes you work more carefully than if you just look for a nice sample text after every change.

Next up

In the next chapter our program doesn't need words, tables or neural layers yet. It only receives pairs of numbers such as “1 becomes 3” and “2 becomes 6”. We examine what could even be called “learning” here.