Many neurons compute with many numbers
You have computed a small network with two inputs and two inner neurons by hand. In a larger network we want to write down the same calculations for many inputs neatly and run them together. For that, all we need at first are lists and tables.
A single number, for example 3, is called a scalar. An ordered list of numbers such as [2,1] is called a vector. A rectangular table of numbers is called a matrix. The umbrella term tensor allows any number of such directions of arrangement. So a tensor doesn't have to be anything mysterious: for us it is an ordered array of numbers with a known shape.
The shape describes the organization, not the content
The list [2,1] contains two numbers and has shape [2]. The table [[1,1],[-1,2]] has two rows and two columns; its shape is [2,2]. The shape doesn't say which values are inside, but how many values are arranged along each axis. An axis here is a direction in the array, such as the row or column direction.
A stack of three tables, each with two rows and four columns, would have shape [3,2,4]. Three numbers in this shape description mean three axes. They don't mean that only three numbers are stored. In total the stack contains 3×2×4=24 numbers.
Our small network as a table
The weights of the two hidden neurons can be written into one shared table:
| Neuron | Weight for first input | Weight for second input | Bias |
|---|---|---|---|
| First | 1 | 1 | −2 |
| Second | −1 | 2 | 1 |
For the inputs [2,1] we compute the familiar neuron sum for each row. First row: 2×1+1×1−2=1. Second row: 2×(−1)+1×2+1=1. Then we apply ReLU to both results.
A matrix multiplication is an organized way to compute many such weighted sums together. It adds no new meaning to our neuron calculation. It just bundles it efficiently.
First a single weighted sum
Take two lists of equal length: [1,2,3] and [4,0,−1]. Multiply the numbers that belong together and add the results: 1×4+2×0+3×(−1)=1. This calculation is called the dot product. The result is a single number, that is, a scalar.
With short names we write:
Here a and b are the two complete lists. The dot between the list names denotes their dot product. Between individual numbers it denotes ordinary multiplication. Which meaning is intended follows from the objects involved.
A large dot product doesn't automatically mean "same meaning." It also depends on how large the components are. We will look at the difference between numerical similarity and meaning later, with learned text vectors.
A matrix multiplication step by step
We want to turn the input list [1,2] into two outputs. The first output should be 1×3+2×5=13. The second should be 1×4+2×6=16. We can organize the required weights like this:
import torch
x = torch.tensor([[1., 2.]])
w = torch.tensor([[3., 4.], [5., 6.]])
print(x @ w) # tensor([[13., 16.]])torch.tensor creates an array of numbers from the lists. The dots in 1. and 2. mark floating-point numbers, which can also represent decimals. In Python the @ sign means matrix multiplication.
For each output column we take its weights, multiply them by the inputs and add up. The shape calculation for this is [1,2] @ [2,2] → [1,2]: an input row with two numbers becomes an output row with two numbers. The inner number 2 has to match, because we need exactly as many weights as there are input components.
Don't confuse it with elementwise multiplication
For two arrays of the same shape, a * b only multiplies the elements in the same positions. It doesn't add them up. a @ b, by contrast, forms weighted sums along a matching axis. Elementwise always means: each position on its own.
A raised T in a formula may also look new: means transpose. Rows and columns are swapped. It is not an exponent. It is needed often, because the same weights are arranged differently depending on the writing convention.
Why PyTorch weights sometimes look flipped
nn.Linear(2,3) denotes a layer with two inputs and three outputs. Each of the three outputs needs two weights. PyTorch stores these weights as three rows of two numbers each, that is, shape [3,2].
from torch import nn
layer = nn.Linear(2, 3)
x = torch.tensor([[2., 1.]])
y = layer(x)
print(layer.weight.shape) # torch.Size([3, 2])
print(y.shape) # torch.Size([1, 3])shape is what PyTorch calls the shape. The layer handles the necessary transposition internally. Here it creates random initial weights, so you shouldn't expect any particular learned prediction.
The shapes that appear later in the language model
A batch is a group of examples processed at the same time. A sequence is an ordered series, later a sequence of text. After a translation step we have yet to explain, each piece of text gets a vector of numbers. For these recurring sizes we use four names:
| Name | Meaning | Small example |
|---|---|---|
| B | Number of examples in the batch | 2 texts |
| T | Number of positions processed per text | 3 positions |
| D | Number of values in each position's vector | 4 numbers |
| V | Number of possible text pieces in the model's vocabulary | for now only a size for later |
The shape [B,T,D]=[2,3,4] then means two texts, three positions each and four numbers per position. It contains 24 numbers. These letters are names for sizes, not operations.
Helper operations in code
Reshape means grouping the same numbers differently, for example twelve numbers into three groups of four. The count of numbers must stay the same. Transpose swaps axes. Broadcasting means a smaller array is reused to fit during an operation, for example the same bias for every input row.
These are handy tools, but you should always know what role each axis plays. In the later attention chapter we will also need a head axis. You don't need to track it from memory yet.
The most important working habit
Before an operation, write down in words what a row and a column mean. Then check the shape. A program can run with matching sizes and still mix up the wrong roles. "Two texts with three positions" is something different from "three texts with two positions," even though both contain six positions.