One input can have several fitting continuations
In the chapter on tokenization you learned that text is split into small building blocks, and each block gets an ID number. These building blocks are called tokens. Our language model should now judge, for the text so far, which token could fit next.
After “The cat sleeps on the”, several continuations are possible. In this example we use only “sofa”, “roof” and “sea”. That is a deliberately tiny vocabulary. A real model has far more options to choose from.
We don't want to store just one single option, but their probabilities. A probability lies between 0 and 1. The value 0.25 corresponds to 25 percent. All the mutually exclusive next options considered here must together have probability 1.
A distribution is a complete split
Suppose the model assigns 60 percent to sofa, 30 percent to roof and 10 percent to sea. These three numbers add up to 100 percent. We call this a probability distribution over the possible next tokens.
It is the model's assessment, not a guarantee about what a person will write. Other texts and other training data can lead to a different assessment. “Sofa is likely here” also does not automatically mean “the complete sentence is factually true”.
The text so far counts too
The probability should depend on the text already read. After “We are taking the boat out to the”, “sea” would probably fit better than after our cat sentence. We call the text already read the context or prefix. Prefix simply means “the part that already comes before”.
There is a short notation for this: . Read it as “probability of sofa, given the text so far”. The P names a probability. The vertical bar means “under the condition”, or more simply here, “after this context”. It is not a division sign.
The network first delivers scores, not percentages
The last layer of the network can output ordinary numbers, such as 2, 1 and 0. Such raw ratings are called scores or, at this output position, logits. They may be negative and don't have to add up to one.
To turn them into probabilities, we use a calculation rule called softmax. It gives higher scores higher probabilities and makes sure that all probabilities together add up to one.
Softmax on three numbers
For the scores [2, 1, 0] we compute in three steps:
- Turn each score into a positive number with the exponential function. That gives roughly [7.389, 2.718, 1].
- Add these numbers: 7.389+2.718+1≈11.107.
- Divide each positive number by this sum.
| Option | Score | Positive conversion | Divided by the sum |
|---|---|---|---|
| sofa | 2 | 7.389 | about 0.665 = 66.5% |
| roof | 1 | 2.718 | about 0.245 = 24.5% |
| sea | 0 | 1 | about 0.090 = 9.0% |
Because of rounding, the displayed percentages can differ very slightly from exactly 100 percent. The unrounded calculation is normalized.
What is the exponential function?
The calculation rule we use is called exp. For this chapter you can treat it like a calculator function: exp(0)=1, exp(1)≈2.718 and exp(2)≈7.389. It can also be written as a power of the number e≈2.718: exp(z)=e to the power of z.
Two properties matter to us: the result is always positive, and larger inputs give larger results. Even a negative score can be turned into a positive number this way. To use softmax, you don't need to master the deeper theory of the number e yet.
Now read the formula
| Part of the formula | In plain words |
|---|---|
| i | Number of the option we are looking at right now |
| Its raw score, for example the score 2 for sofa | |
| The positive conversion of this score | |
| j in the summation sign | Running number through all options |
| Sum of all positively converted scores | |
| The probability we are looking for, for this one option |
So the formula only says: “Positive conversion of the option we are looking at, divided by the sum over all options.” For sofa that is 7.389/11.107≈0.665. The i in the numerator stays the one selected option; the j in the denominator runs through the whole list.
Why the other probabilities move too
If you raise only the sofa score, its positive value grows. At the same time, the sum in the denominator grows. That is why the probabilities of roof and sea go down, even though their raw scores stay the same. The options share 100 percent between them.
This explains why softmax does not compute independent yes/no probabilities. Here we choose exactly one next token from a shared set of options.
Several tokens in a row
Imagine a sequence of only two tokens. The first gets probability 0.5. The second gets probability 0.2 after the first. The joint probability of this particular two-token sequence is 0.5×0.2=0.1.
For longer sequences we likewise multiply the probability of each next token given its respective preceding context. This compact notation stands for that:
Read step by step: is the token at position t. are all tokens before it. T is the length of the sequence. The product sign says that we multiply the individual conditional probabilities from position 1 to T. The dots on the left mean “all elements in between”.
This decomposition is called the chain rule of probability. It shares the idea of putting things together with other chain rules, but it is not the derivative rule from backpropagation. For the language model, it turns “judge a whole text” into many tasks of the form “judge the next token”.
The experiment and later extensions
In the experiment you can change the three scores. The temperature slider changes how strongly softmax emphasizes the differences: smaller positive values lead to a more concentrated distribution, larger ones to a more even one. To do this, each score is divided by the temperature value before softmax. We explain the final choice of a token in detail in the chapter on generation.
The displayed entropy is a measure of how spread out or concentrated the probabilities are. You don't need to know its formula now. When almost all the probability sits on one option, it is small; with an even split it is larger.
For very large raw scores, the code uses a numerically stable variant: it first subtracts the largest score from all scores. This shared offset does not change softmax, but it prevents needlessly huge intermediate numbers. This implementation detail changes nothing about the three calculation steps you just learned.