Tokenwerk · The LLM textbook

Chapter 5 · I · Understanding neural networks · 5 minutes

From neurons to a small network

Two inputs, two inner neurons and one output. Every connection gets a numerical value you can follow.

One result becomes the next input

A neural network is formed when we connect computing units. The output of one neuron can serve as the input to further neurons. We group several neurons that compute at the same stage into a layer.

Our small network has two input values. From these, two neurons compute two intermediate values. A final neuron combines these intermediate values into one output. We need no words and no large matrices to understand this process.

The names of the layers

The input layer here refers to the numbers that are provided. No trainable calculation happens yet at these input nodes. The hidden layer lies between input and output and computes intermediate values. The output layer produces the prediction that we compare with the target.

“Hidden” doesn't mean the numbers are inaccessible. We can print them and check them. What it means is: the dataset normally doesn't specify its own targets for these intermediate values. We see the input and the desired final result, but no prescribed answers for each inner unit.

Our shared numerical example

The two inputs are again 2 and 1. The hidden layer contains two ReLU neurons. Their weights differ, even though both see the same inputs.

Neuron First contribution Second contribution Bias Sum After ReLU
First inner neuron 2 × 1 = 2 1 × 1 = 1 −2 1 1
Second inner neuron 2 × (−1) = −2 1 × 2 = 2 +1 1 1

Here both intermediate values happen to equal one. They still come from different calculations. If we change an input, they don't have to stay equal.

The output neuron multiplies the first intermediate value by 2 and the second by 3. Its bias is zero. In this example it uses no final ReLU:

Prediction = first intermediate value × 2 + second intermediate value × 3 = 1 × 2 + 1 × 3 = 5.

With that, we have run the entire network by hand.

What “forward” means

We call this path from the inputs to the computed output the forward pass. In such a pass, only the fixed calculations are carried out at first. The weights don't change during it.

In the experiment you can change the inputs and watch how both intermediate values and the output change. No training happens there. You are using the same blueprint with the same weights for new inputs.

This distinction matters: “the result changes” does not automatically mean “the network is learning”. A calculator also gives different results for different inputs without changing its calculation rules.

How many parameters does this network have?

Each of the two inner neurons has two weights and one bias. That is three parameters each, six in total. The output neuron also has two weights and one bias, so another three. In total our network has nine parameters.

The two inputs don't count. The two computed intermediate values aren't permanent parameters either. They are computed anew for each example. These intermediate values are called activations or internal states.

Later we'll talk about millions of parameters. The counting idea is the same: we count the adjustable settings, not the number of words read or the number of possible inputs.

Width and depth

Width roughly means how many units or numbers there are per stage. Our hidden layer has width two. Depth describes the number of successive processing layers. Whether the input or output layer is included in such a figure isn't consistent everywhere; so wherever possible, we name the concrete trainable layers.

A wider network can carry out more different intermediate calculations in parallel. A deeper network can process results further across more stages. Neither automatically increases quality. The data and the training have to suit this blueprint.

What do the inner neurons mean?

We haven't assigned the inner neurons any meanings such as “grammar” or “knowledge”. Here they are, for now, just two calculation paths. Useful patterns can develop during learning. But there is no guarantee that each individual unit has a human property that is easy to name.

The capabilities come from the interplay of the network. That is one reason why it can later be difficult to explain large networks just by looking at individual weight values.

The next question

Suppose the correct target for our example were 6. Our network gives 5. So we know the prediction isn't right. But how do we turn that into a learning signal? After a short introduction to notation, we define an error measure for this: a number that says how far the model is from the desired result.