Tokenwerk · The LLM textbook

Chapter 13 · II · From numbers to language · 6 minutes

Language as a probability problem

How continuing a text becomes a precise learning task — and what probabilities do and don’t say about truth.

One input can have several fitting continuations

In the chapter on tokenization you learned that text is split into small building blocks, and each block gets an ID number. These building blocks are called tokens. Our language model should now judge, for the text so far, which token could fit next.

After “The cat sleeps on the”, several continuations are possible. In this example we use only “sofa”, “roof” and “sea”. That is a deliberately tiny vocabulary. A real model has far more options to choose from.

We don't want to store just one single option, but their probabilities. A probability lies between 0 and 1. The value 0.25 corresponds to 25 percent. All the mutually exclusive next options considered here must together have probability 1.

A distribution is a complete split

Suppose the model assigns 60 percent to sofa, 30 percent to roof and 10 percent to sea. These three numbers add up to 100 percent. We call this a probability distribution over the possible next tokens.

It is the model's assessment, not a guarantee about what a person will write. Other texts and other training data can lead to a different assessment. “Sofa is likely here” also does not automatically mean “the complete sentence is factually true”.

The text so far counts too

The probability should depend on the text already read. After “We are taking the boat out to the”, “sea” would probably fit better than after our cat sentence. We call the text already read the context or prefix. Prefix simply means “the part that already comes before”.

There is a short notation for this: P(sofa∣text so far)P(\text{sofa}\mid\text{text so far}). Read it as “probability of sofa, given the text so far”. The P names a probability. The vertical bar means “under the condition”, or more simply here, “after this context”. It is not a division sign.

The network first delivers scores, not percentages

The last layer of the network can output ordinary numbers, such as 2, 1 and 0. Such raw ratings are called scores or, at this output position, logits. They may be negative and don't have to add up to one.

To turn them into probabilities, we use a calculation rule called softmax. It gives higher scores higher probabilities and makes sure that all probabilities together add up to one.

Softmax on three numbers

For the scores [2, 1, 0] we compute in three steps:

  1. Turn each score into a positive number with the exponential function. That gives roughly [7.389, 2.718, 1].
  2. Add these numbers: 7.389+2.718+1≈11.107.
  3. Divide each positive number by this sum.
Option Score Positive conversion Divided by the sum
sofa 2 7.389 about 0.665 = 66.5%
roof 1 2.718 about 0.245 = 24.5%
sea 0 1 about 0.090 = 9.0%

Because of rounding, the displayed percentages can differ very slightly from exactly 100 percent. The unrounded calculation is normalized.

What is the exponential function?

The calculation rule we use is called exp. For this chapter you can treat it like a calculator function: exp(0)=1, exp(1)≈2.718 and exp(2)≈7.389. It can also be written as a power of the number e≈2.718: exp(z)=e to the power of z.

Two properties matter to us: the result is always positive, and larger inputs give larger results. Even a negative score can be turned into a positive number this way. To use softmax, you don't need to master the deeper theory of the number e yet.

Now read the formula

pi=exp⁡(zi)∑jexp⁡(zj).p_i=\frac{\exp(z_i)}{\sum_j\exp(z_j)}.
Part of the formula In plain words
i Number of the option we are looking at right now
ziz_i Its raw score, for example the score 2 for sofa
exp⁡(zi)\exp(z_i) The positive conversion of this score
j in the summation sign Running number through all options
∑jexp⁡(zj)\sum_j\exp(z_j) Sum of all positively converted scores
pip_i The probability we are looking for, for this one option

So the formula only says: “Positive conversion of the option we are looking at, divided by the sum over all options.” For sofa that is 7.389/11.107≈0.665. The i in the numerator stays the one selected option; the j in the denominator runs through the whole list.

Why the other probabilities move too

If you raise only the sofa score, its positive value grows. At the same time, the sum in the denominator grows. That is why the probabilities of roof and sea go down, even though their raw scores stay the same. The options share 100 percent between them.

This explains why softmax does not compute independent yes/no probabilities. Here we choose exactly one next token from a shared set of options.

Several tokens in a row

Imagine a sequence of only two tokens. The first gets probability 0.5. The second gets probability 0.2 after the first. The joint probability of this particular two-token sequence is 0.5×0.2=0.1.

For longer sequences we likewise multiply the probability of each next token given its respective preceding context. This compact notation stands for that:

P(x1,…,xT)=∏t=1TP(xt∣x<t).P(x_1,\ldots,x_T)=\prod_{t=1}^{T}P(x_t\mid x_{<t}).

Read step by step: xtx_t is the token at position t. x<tx_{<t} are all tokens before it. T is the length of the sequence. The product sign says that we multiply the individual conditional probabilities from position 1 to T. The dots on the left mean “all elements in between”.

This decomposition is called the chain rule of probability. It shares the idea of putting things together with other chain rules, but it is not the derivative rule from backpropagation. For the language model, it turns “judge a whole text” into many tasks of the form “judge the next token”.

The experiment and later extensions

In the experiment you can change the three scores. The temperature slider changes how strongly softmax emphasizes the differences: smaller positive values lead to a more concentrated distribution, larger ones to a more even one. To do this, each score is divided by the temperature value before softmax. We explain the final choice of a token in detail in the chapter on generation.

The displayed entropy is a measure of how spread out or concentrated the probabilities are. You don't need to know its formula now. When almost all the probability sits on one option, it is small; with an even split it is larger.

For very large raw scores, the code uses a numerically stable variant: it first subtracts the largest score from all scores. This shared offset does not change softmax, but it prevents needlessly huge intermediate numbers. This implementation detail changes nothing about the three calculation steps you just learned.