Tokenwerk · The LLM textbook

Chapter 7 · I · Understanding neural networks · 5 minutes

How do we measure a wrong answer?

Before a network can learn, we need a number for how big its error is. We start with a prediction of 2 instead of 6.

Right or wrong is often too coarse

Our small model should output 6 for input 2. With its current setting, though, it outputs 2. A prediction of 5 would also be wrong, but it would be much closer to the target. We want to be able to measure that improvement.

For this we define an error function, also called a loss function. It takes the prediction and the target and returns a number. In our example: the smaller this number, the better the prediction fits.

First compute the difference

We subtract the target from the prediction. For prediction 2 and target 6 this gives 2−6=−4. The negative sign tells us the prediction is too low. For prediction 10 we get 10−6=4, so the prediction is too high.

Both errors have the same size. If we simply added up the signed differences of many examples, they could cancel out: −4+4=0. A wrong calculation would then look error-free by accident.

Square the difference

A simple fix is to multiply each difference by itself. Both (−4)² and 4² give 16. An exact prediction has difference zero and therefore error zero.

Prediction Target Difference Squared error
2 6 −4 16
4 6 −2 4
5 6 −1 1
6 6 0 0
7 6 1 1
10 6 4 16

"Squared" just means we use the square of the difference. As a result, large deviations count more heavily than small ones.

The formula describes exactly this table

We call the prediction y^\hat y, pronounced "y-hat", and the target yy. We call the error number LL, for loss.

L=(y^−y)2.L=(\hat y-y)^2.

Read the formula: Subtract the target from the prediction. Multiply the difference by itself. Call the result L.

The hat is only a label marking the prediction. The 2 at the top is an exponent. With the numbers 2 and 6 you get (2−6)²=16 again. So you don't need to learn a new calculation to understand this formula.

An error value doesn't yet tell us how to learn

If we know that L=16, we know how big this error is. That alone doesn't tell us which of nine, or millions of, parameters should be changed, and how.

With the one-parameter model we can try out different settings. With very many parameters, searching by trying each one separately would be costly. So in the next chapter we learn to find the direction of change with a derivative. The error function is the goal; the update rule is the way to improve on that goal.

Looking at several examples together

Take our model "input times weight" with weight 2. The training examples are 1→3, 2→6 and 3→9. It predicts 2, 4 and 6. The differences are −1, −2 and −3. Their squares are 1, 4 and 9.

The mean squared error is (1+4+9)/3=14/3≈4.667. "Mean" means we add up the errors and divide by the number of examples. This common error value is called the mean squared error, or MSE for short.

The average makes runs with similarly structured examples easier to compare than an ever-growing sum would. It can still hide individual large errors. For a thorough evaluation, we will later also look at concrete tasks.

Why not just take the absolute value?

The absolute difference would also be a possible error measure. For 2 instead of 6 it is 4. The squared error, by contrast, is 16. Different error measures weigh large and small deviations differently and can lead to different learning behavior.

So there is no single "error" for every task. For our number prediction, the squared error is easy to understand. For a language model, by contrast, we want to score a distribution over possible next pieces of text. For that we will introduce a different error measure later. There, too, we start with numeric examples.

What a small error value shows

A small error value shows that the predictions on the measured examples fit the targets used. It does not prove that unseen examples will work too. Likewise, a model can imitate a faulty target very well.

Training data is used to change the weights. Validation data consists of held-back examples that we use along the way to check how the model does on data not used for learning. We will cover a final independent test in more detail later.

In the experiment

Change only the prediction and watch the difference and its square. Then set two predictions that lie equally far to the left and right of the target. The error value should be the same. Its size and the direction of the needed improvement are two different pieces of information.