Tokenwerk · The LLM textbook

Chapter 17 · III · Building the transformer · 6 minutes

Self-attention without the mystery

One position mixes information from other positions. We first compute three contributions by hand and only then read the attention formula.

What our position vectors are still missing

An embedding is the learned list of numbers for a token ID. At first, the same token gets the same table vector, even when it appears in different sentences. We need a computation step that lets one text position take in information from other positions.

This step is called self-attention. "Self" means the information comes from the same input sequence. "Attention" here describes a computed weighting. It is neither conscious paying attention nor a ready-made explanation of how the model thinks.

We start with exactly one reading position and three available sources of information. The question is: how strongly should each source flow into the new list of numbers?

First, a weighted mix

Picture three values: 10, 20 and 30. In this context, value means "the content to be passed along". To keep things simple, each piece of content here is just one number. In a real model it is a list of numbers.

With weights 0.5, 0.25 and 0.25, the mix is: 0.5×10+0.25×20+0.25×30=17.5. A larger weight takes more from the corresponding source. The weights add up to one.

You already know this calculation as a weighted sum. What is new is that the weights are computed from the current text instead of being fixed connection weights that are trained directly.

Three roles for the same text position

From the internal vector of each position, learned linear layers produce three different lists of numbers:

Name Translation and role
Query, Q for short "Search request": which pattern does the reading position compare the sources against?
Key, K for short "Comparison feature": which list of numbers does a source offer for comparison?
Value, V for short "Content": which list of numbers is taken over in the mix?

The search analogy only explains the roles. The model does not write human search terms and does not look anything up in an external database. It computes lists of numbers through multiplications and sums. Projection is the technical name for such a learned linear conversion into a new vector space; for now, "compute a different list of numbers" is all you need.

Comparing query and key

Our query is [1,0]. The three keys are [1,0], [0,1] and [1,1]. We compare the query with each key using a dot product: multiply matching components and add them up.

Source Calculation Comparison score
Key 1 = [1,0] 1×1 + 0×0 1
Key 2 = [0,1] 1×0 + 0×1 0
Key 3 = [1,1] 1×1 + 0×1 1

Here the first and third sources get the same score. That only tells us something about these chosen numbers. It does not prove any semantic similarity between real words.

Turning scores into mixing weights

Our query and keys each have two components. We call this length d=2. We divide the scores by the square root of d, which is about 1.414. This turns [1,0,1] into roughly [0.707; 0; 0.707].

Under certain typical starting assumptions, this scaling keeps the scores at a more manageable size as vectors grow. For the first calculation, this is enough: it is a fixed size adjustment before softmax, not a newly learned parameter.

Now we apply the softmax you already know. The weights come out at roughly [0.401; 0.198; 0.401]. They are positive and add up to one. So our position takes about 40.1 percent from value 1, 19.8 percent from value 2 and 40.1 percent from value 3.

Computing the new output

With values 10, 20 and 30 we get approximately:

0.401×10 + 0.198×20 + 0.401×30 = 20.

That the result is exactly 20 comes from the symmetric choice of this example. If you change only value 3 in the experiment, the weights stay the same, but the output changes. If you change the query instead, the weights change. This way you can keep the comparison and the content being passed along apart.

Computing many positions at once

For a whole text, we compute one query per position. Each query gets its own row of comparison scores against the keys. The resulting table is called the score matrix. A row means "who is reading", a column means "who is being read from".

With three text positions, the table has three rows and three columns. We apply softmax to each row separately. So we do not normalize all nine numbers together, but each reading position's choice on its own.

The familiar calculation in short form

A=softmax⁡(QK⊤d),Y=AV.A=\operatorname{softmax}\left(\frac{QK^\top}{\sqrt d}\right),\qquad Y=AV.
Symbol Meaning in the calculation just explained
Q Table of all query vectors
K Table of all key vectors
K⊤K^\top The same key table with rows and columns swapped
QK⊤QK^\top All query–key dot products as one shared table
d Number of components per query or key
A Table of mixing weights after row-wise softmax
V Table of the contents to be mixed, i.e. the values
Y Table of the new outputs, one per reading position

Read the formula: compare queries with keys, scale the scores, compute softmax per row, and use these weights to mix the values. The comma in the formula separates two consecutive calculation steps.

Note the name clash: in this section, capital V means the value table. In shape descriptions like [B,T,V], however, V stands for the vocabulary size. This kind of reuse is common in technical texts, but confusing. We will always state the role explicitly.

What is trained and what is recomputed

The weights of the linear layers that produce Q, K and V are model parameters. They change during training. The attention weights A, on the other hand, are intermediate values computed fresh for every input. They are not stored in the model as one fixed table for all texts.

Backpropagation can also compute through the dot products, the softmax and the mixing. This is how the projections learn which comparisons are useful for the actual next-token objective. There is no human-specified attention plan for each position.

An open question leads to the next chapter

So far, every source was allowed to take part in the mix. But when predicting the next token, a position must not see a future answer. We have to block certain cells of the table. This restriction is called a mask. The next chapter builds it on top of the score matrix we just explained.