Why we use formulas at all
In the neuron chapter we wrote “first input times first weight, plus second input times second weight, plus bias”. That is understandable, but with ten inputs it gets long. A formula abbreviates these instructions. It is meant to express a familiar calculation more compactly, not to replace an explanation.
You don't have to memorize this chapter. Use it as a dictionary and come back via the chapter links when you need to. We still explain new symbols where they are needed, too.
A letter is a name for a number
When we say “x is 2”, we give the number a name. A variable name such as x can stand for a different number in another example. It is not a multiplication sign. For multiplication we use × in detailed examples, and in formulas often a dot or letters written directly next to each other.
3 × x, and mean the same calculation in this context. With x = 2, the result is 6. The equals sign = says that the left and right sides have the same value. The sign ≈ means “approximately equal”, for example because of rounding.
A function is a calculation rule with a name
If we write , this means: the rule f takes a number x and multiplies it by three. means “apply this rule to 2”. The result is 6.
So the parentheses in f(2) are not an extra multiplication. Compare it with a function in a program: triple(2) returns 6. The mathematical function name f is just shorter.
For a nested calculation such as , you work from the inside out: first apply f to 2, then apply g to the result. This is similar to how we connect the layers of a network.
Small numbers at the top and bottom mean different things
| Notation | Meaning | Example |
|---|---|---|
| The second element of a collection | In [4,7,9] the second element is 7 | |
| Multiply x by itself | With x=3, x²=9 | |
| The element one place after position t | Position 4 comes after position 3 | |
| All elements before position t | Before position 3 come positions 1 and 2 |
The small number at the bottom is called an index. It labels a position. In mathematical examples we often count from one; Python counts list positions from zero. So for code, we say explicitly which index is meant.
A hat over a letter, as in , is read “y-hat”. In this book it usually stands for a prediction. without a hat then denotes the desired target. The hat is not a calculation; it distinguishes two roles.
A sum is repeated addition
Take the list [2,4,6]. Its sum is 2+4+6=12. The general shorthand is:
Read the formula: Add up the list elements while the running number i goes from 1 to 3. The large symbol Σ is called sigma and stands for a sum. The start is written below the symbol, the end above it.
The i is a running variable (an index variable): a placeholder for the numbers plugged in one after another. If we called it j, it would be the same calculation. What matters is using the same name in all the places that belong together.
A product is repeated multiplication
For the same list, the product is 2×4×6=48. The shorthand uses a large Π, read as “product sign”:
So a sum and a product are different instructions. Later we'll meet products of probabilities. Then we'll first write down two or three concrete factors and only afterwards this shorthand.
Fractions, the mean and the square root
A fraction bar means division. “Sum divided by count” gives the arithmetic mean. For 2, 4 and 6 that is (2+4+6)/3=4.
The square root of 9 is 3, because 3×3=9. The square root of 2 is about 1.414. Its symbol is √. In this book, square roots are used, among other things, to adjust the scale of numbers.
An absolute value is a number's distance from zero: both 4 and −4 have absolute value 4. So the absolute value says nothing about the sign. This idea later helps to distinguish the size of an error from its direction.
Greek letters are just names too
For a learning rate we'll later meet η, pronounced “eta”. A very small helper value is often called ε, “epsilon”. Such letters don't have any special computational effect by themselves. They are agreed names whose meaning we have to define.
The same symbol can mean something different in different technical texts. That is why a good formula comes with an explanation of its symbols. In this textbook you'll find that explanation right next to the calculation in question.
What comes later
We explain derivatives using a small change in a weight in the chapter on learning. We explain the exponential function and the logarithm when converting scores into probabilities and in the language model's error measure, respectively. You don't need to know these tools yet.
Our way of working stays the same: first a question, then numbers, then a short notation. If a formula looks new, try replacing its letters with the numbers from the example. Often that makes it readable right away.