Tokenwerk · The LLM textbook

Chapter 2 · I · Understanding neural networks · 4 minutes

What does it mean for a program to learn?

One input, one desired output and one adjustable number. That is enough to see the whole learning process.

First, an ordinary calculation rule

Imagine we want to turn the number 2 into the number 6. If we know the rule, we simply program: “Multiply the input by three.” For the input 4, the program then returns 12.

For many tasks, though, we don't know the rule completely. How exactly do you recognize a handwritten letter? Which words fit at the end of a sentence? You can show examples, but it is hard to write down every necessary rule by hand. Machine learning tries to tune a suitable calculation rule using such examples.

To get started, we still use a simple rule. That way we can always check whether the method is working sensibly.

We give examples instead of the finished setting

Our program receives this table:

Input Desired output
1 3
2 6
3 9

One row is called a training example. The input is what the model gets to see. The desired output is called the target. We call the whole collection of such examples the training data.

We allow the model a very simple calculation: input times an adjustable number. This adjustable number is a parameter. It starts at, say, 1. Then for the input 2 the model first predicts 2. But the target is 6.

What we specify and what is learned

We specify the blueprint “input times parameter”. In this example the model invents neither new lines of code nor a different kind of arithmetic. What changes during learning is the parameter value.

With parameter 2, the input 2 becomes the prediction 4. With parameter 3, it becomes 6. We can try out which setting fits our examples better. Later this adjusting is done automatically.

A model here is the chosen calculation rule together with its current parameters. Training means improving the parameters based on examples. We only cover the rule for changing the parameters in the learning chapter. For now it is enough that something can be adjusted at all.

The prediction is not the target

The target comes from the example. The model computes the prediction itself. We have to keep the two apart, otherwise we can't detect errors.

Term In our example
Input 2
Current parameter 1
Computed prediction 2 × 1 = 2
Desired target 6
Difference Prediction is 4 too low

The word “prediction” doesn't imply the future. Even when all the numbers are already known, we call the output computed by the model a prediction.

What happens after training?

Once the parameter is set, we want to use an input that wasn't in the table: 4. The learned parameter 3 then gives 12. Applying the learned model is called inference. For now you can simply read this word as “using the model”.

When a model works well on new, suitable examples, we speak of generalization. It shouldn't just reproduce the examples it was shown, but carry over a useful pattern. Our simple model can't store arbitrary answers: a single multiplication factor has to fit all rows.

Why not just store every answer?

A lookup table could return the three known rows exactly. But for the new input 4 it has no stored answer. A well-chosen model can express relationships between the examples through its computational structure.

This is also a limitation. If the desired outputs follow a completely different rule, our single multiplication factor may not be enough. Then there is no setting that describes all examples correctly. That is not a flaw in the search method but a blueprint that is too simple.

For example, “input times a number” can only ever return 0 for input 0. If the output should already be 5 at 0, we additionally need a shift. That is exactly why we meet the bias in the next chapter.

For the language model later on

For a language model, the examples aren't just small pairs of numbers. It sees the text so far and should compute a suitable continuation. The basic roles still stay the same: input, desired output, computed output and adjustable parameters.

There are far more parameters and a more complex blueprint. We build this blueprint up step by step. The idea of learning you have met here remains recognizable throughout.