We improve the same small calculating machine
You now know weights, predictions and an error value. Our model multiplies an input by a weight. The input is 2, the desired target 6 and the current weight 1. So the prediction is 2 and the squared error is 16.
How should we change the weight? We could start by trying two nearby settings. With weight 0.99 the model outputs 1.98 and an error of 16.1604. With weight 1.01 it outputs 2.02 and an error of 15.8404. Slightly larger is clearly better here.
A slope from two measurement points
The weights 0.99 and 1.01 are 0.02 apart. Between them, the error changes from 16.1604 to 15.8404, that is, by −0.32. Dividing the change in error by the change in weight gives −0.32/0.02=−16.
This number describes the slope of the error curve near the current weight. A negative result says: if the weight grows a little, the error falls. A positive result would mean the error rises as the weight gets larger.
A derivative describes this local slope in the limit of ever smaller changes. "Local" means near the setting we are currently looking at. It says nothing about what a huge jump somewhere else would do.
What the tricky symbol means
The error is called , the weight . The notation for the derivative is:
Read this as: "the derivative of the error L with respect to the weight w." In plain terms we are asking: how sensitively does the error react to a small change in this weight? In this example the answer at the start is −16.
The symbol is fixed notation for a derivative. You don't have to read it as an ordinary fraction of two numbers you already know. The calculation with two measurement points is an approximation you can use to check what it means.
Reverse the direction
The derivative shows how the error rises or falls as the weight increases. We want to lower it. So we change the weight against the slope.
With slope −16 that means: increase the weight. With slope +16 it would mean: decrease the weight. We set the step size with a learning rate. In our example, a learning rate of 0.05 means we use 0.05 times the derivative as the correction.
The calculation is: new weight = old weight − learning rate × derivative. With our numbers we get 1−0.05×(−16)=1.8. A minus in front of a negative number leads to an increase here.
Only then the short form
| Symbol | Meaning here | Value |
|---|---|---|
| Weight before the learning step | 1 | |
| Weight after the learning step | 1.8 | |
| Learning rate; "eta" is just its short name | 0.05 | |
| Slope of the error at the current weight | −16 |
This learning rule is called gradient descent: we move in the opposite direction of the gradient, that is, downhill. With only one weight, one derivative is enough. With many weights we collect one such derivative for each weight. This collection is called the gradient.
We check the step instead of trusting it blindly
With weight 1.8 the prediction becomes 1.8×2=3.6. The new difference from the target is 3.6−6=−2.4. Its square is 5.76. The error really is smaller than the previous 16.
It isn't the right answer yet. Learning usually consists of many such steps. Before each step we compute the derivative at the new setting. We don't keep using the initial −16 forever.
Where does the computer get the derivative?
For our simple example you can derive a direct rule: derivative = 2 × difference × input. The factor 2 comes from the square in our error function. The factor "input" appears because a change in the weight first changes the prediction by exactly that multiple.
With difference −4 and input 2 we get 2×(−4)×2=−16. In the next chapter we explain this path through several calculation steps in more detail. You don't need to be able to derive the rule from the algebra yourself yet.
Large networks use a method that systematically combines such local rules of change. We will later meet the automatic differentiation it relies on as autograd.
Why big steps can fail
In the experiment, choose a learning rate of 0.3. The first step gives 1−0.3×(−16)=5.8. The prediction is then 11.6 and the error (11.6−6)²=31.36. Even though the initial direction was right, the step is so large that it overshoots the good region.
That's why a local derivative doesn't say "go as far as you like in this direction." The learning rate belongs to the learning method. It is a setting we choose; in this example it is not itself a weight of the model.
What you should understand now
A learning step has three separate jobs: compute a prediction, determine the local effect of the weights on the error, and change the weights carefully. The first job is the forward pass. In larger networks the second is done by the backward pass. The third is handled by an update rule, often called an optimizer.
In the next chapter each of these jobs stays visible. We take two calculation steps in a row and follow how an early weight still contributes to the later error.