Tokenwerk · The LLM textbook

Chapter 4 · I · Understanding neural networks · 4 minutes

Why neurons need an activation function

A simple rule after the sum makes an important difference. We start with: negative values become zero.

We add a rule after the sum

In the last chapter, the neuron computed the number 1.5 from its inputs. Now we decide how this sum is passed on. An activation function is a fixed calculation rule that is applied to the sum.

Our first rule is: if the number is negative, output zero. Otherwise, pass the number on unchanged. This rule is called ReLU, pronounced roughly “ray-loo”. The abbreviation stands for Rectified Linear Unit. You don't need to memorize the full name; what matters is the effect.

Sum before the rule Output after ReLU
−3 0
−0.5 0
0 0
1.5 1.5
4 4

The number after the activation function is also called the activation. So “activation function” names the rule, while “activation” often means the concrete result of that rule.

The short notation

zz again stands for the sum before the activation function. We call the new output aa. That is just a name; it is not a new parameter.

a=ReLU⁡(z)=max⁡(0,z).a=\operatorname{ReLU}(z)=\max(0,z).

Read the formula: The activation a is the ReLU output for z. To get it, pick the larger of the two numbers zero and z. max means “maximum”, that is, the larger value. ReLU and max describe the same rule here.

For our neuron with sum 1.5, the output is 1.5. If we change its weights so that the sum is −2, the output becomes zero. In the experiment you can see this transition directly.

Why not just pass every sum on unchanged?

Consider a chain of two simple calculation rules. The first multiplies the input by 2. The second multiplies the result by 3. Overall, the input is multiplied by 6. The two steps can be combined into a single one.

Adding shifts doesn't solve this problem either. First “times 2, plus 1” and then “times 3, plus 4” gives “times 6, plus 7” overall. Despite two stages, the kind of calculation has not become any richer.

If you plot input and output in a coordinate system, such rules draw a straight line. With a shift, they are, strictly speaking, called affine in mathematics. When describing neural networks, the corresponding layers are nevertheless often called “linear”. We always say whether a bias is included.

A kink changes what is possible

ReLU does not draw a single straight line over the whole range of numbers. To the left of zero the output stays at zero; to the right it rises. At the boundary there is a kink. This rule is nonlinear: it can't be expressed as a single multiplication plus shift for all inputs.

By combining several such rules, a network can respond differently to different ranges of inputs. For example, it can take a change into account only once a certain value is exceeded.

Still, a single ReLU neuron can't solve every possible task. The greater expressive power comes from the interplay of several units, their weights and their activation functions. In the next chapter we work through such a small network completely.

The neuron is not consciously on or off

The vivid phrase “the neuron fires” can be misleading. Our program simply computes numbers. A ReLU output of zero means that this particular output is zero in this calculation. An output of 3 is not a biological impulse and not a human decision.

“Nonlinear” isn't some kind of magic either. It describes a property of the function being computed. You can check the function point by point in the experiment.

Why other names show up later

There are various activation functions. Some change negative values gently instead of setting them sharply to zero. When we use GELU or SiLU later, we explain their role again at that point. For getting started, ReLU is enough: you can compute every output without a calculator.

The last layer of a network doesn't have to use the same activation function as the inner layers. If it should produce arbitrary positive and negative scores, an unchanged weighted sum can be suitable. If it should deliver probabilities, we need a different conversion, which we meet in the language model part.

Predict, then try it out

First think about what happens when you shift the bias. A larger bias can push the sum from the negative into the positive range. Then a ReLU output that was zero before becomes positive. The activation rule itself was not changed; only its input has shifted.