Tokenwerk · The LLM textbook

Chapter 3 · I · Understanding neural networks · 4 minutes

An artificial neuron you can compute by hand

At first, a neuron is a small computing unit: weight the inputs, add up the contributions and add a baseline shift.

What is an artificial neuron?

An artificial neuron takes in numbers and computes a new number from them. The name is inspired by biological nerve cells. For our calculation, though, you don't need to know any biology. An artificial neuron is a heavily simplified mathematical computing unit, not a full simulation of a nerve cell.

We start with two inputs. Each input gets its own adjustable multiplication factor, a weight. The neuron multiplies each input by its weight and adds up the results. Then one more adjustable number is added: the bias, a baseline shift.

Many artificial neurons then apply an additional calculation rule to the result. We cover this activation function in the next chapter. Here we focus on the weighted sum first.

An example with two inputs

We choose the inputs 2 and 1. The first weight is 1, the second −1. The bias is 0.5. These are deliberately small numbers with no hidden meaning; we want to see the calculation in full.

Step Calculation Result
First contribution First input × first weight 2 × 1 = 2
Second contribution Second input × second weight 1 × (−1) = −1
Add the contributions 2 + (−1) 1
Add the bias 1 + 0.5 1.5

So the weighted sum is 1.5. A negative weight means that a larger positive input pulls the sum down. It does not mean “bad weight”. Which sign is useful depends on the task.

Why weights matter

Keep the inputs the same and change only the first weight from 1 to 2. The first contribution grows from 2 to 4. The total then becomes 4 − 1 + 0.5 = 3.5.

With this setting, the neuron responds more strongly to the first input. With weight 0, that input has no direct influence on the sum. With weight −2, it acts with the opposite sign.

So you can remember a “weight” as the strength and direction of an influence. The weights are set later during training. The inputs, by contrast, come from the particular example.

Why add a bias?

Set both inputs to zero. Without a bias, the sum would always be zero, whatever the weights. With bias 0.5, it starts at 0.5. The bias shifts the starting point.

A simple everyday example is calculating a price: there is a base price and a price per item. The price per item acts like a weight, the base price like a bias. This analogy only explains the calculation. An artificial neuron doesn't actually have to compute prices.

Weights and bias are both parameters, that is, trainable settings. An input is not a parameter. If you give a fully trained network a different text, its inputs change; the learned weights initially stay the same.

The same calculation with short names

So that we don't have to keep writing out “first input”, we give the numbers names:

Symbol Read it as Value here
x1x_1 first input 2
x2x_2 second input 1
w1w_1 weight of the first input 1
w2w_2 weight of the second input −1
bb bias, the baseline shift 0.5
zz result before the activation function 1.5

The small subscript number is a label, not a power. x2x_2 means “the second input”, not “x squared”.

Now we can write down exactly the familiar calculation in short form:

z=x1w1+x2w2+b.z=x_1w_1+x_2w_2+b.

Read the formula: Multiply the first input by its weight. Add the second input times its weight. Add the bias. Letters written next to each other mean multiplication here.

If we plug in our numbers, we again get 2 × 1 + 1 × (−1) + 0.5 = 1.5. The formula adds no new idea; it only shortens the description.

One neuron and many examples

The same neuron is used for different inputs, one after another. Its weights and its bias are shared across all examples. Training tries to find a setting that is useful across many examples.

Later a network has a great many neurons and parameters. Even so, this single operation remains an important building block. If you can work through this example, you already know the basic operation of a large class of neural layers.