We need an error measure for the next piece of text
In the small number model, we subtracted the prediction and the target from each other. In the language model, the output is a whole list of probabilities. The training example tells us which token actually came next. We want to reward the model for giving this token a high probability.
Suppose that in the training document “The cat sits” is followed by the token “on”. If the model gives “on” only 10 percent, the error should be larger than at 80 percent. The rule must still acknowledge that language can have more than one plausible continuation. For now, our training target remains the observed continuation.
Why the ID number itself is not a target value
If “on” has ID 42 and “under” ID 43, the numerical difference of one is meaningless. The IDs are addresses in the tokenizer. “Under” would not automatically be almost right just because its ID is numerically close to 42.
So we don't use the squared distance between token IDs. Instead we assess the probability the model assigns to the correct token. Whether its ID is 42 or 900 makes no difference to this principle.
Matching inputs to targets
Take the sequence [BOS, The, cat, sleeps, EOS]. BOS is the start marker you have already met, EOS the end marker. This sequence gives four learning tasks:
| Visible input so far | Desired next token |
|---|---|
| BOS | The |
| BOS, The | cat |
| BOS, The, cat | sleeps |
| BOS, The, cat, sleeps | EOS |
In the program we store the input [BOS, The, cat, sleeps] and the target [The, cat, sleeps, EOS]. The two lists are shifted against each other by one position. This is called a target shift. The causal mask later will prevent any computation path from already reading its target, which stands to the right.
A suitable error curve
Our error rule should give zero at probability 1. For small correct probabilities it should grow sharply. For this we use the negative natural logarithm:
| Probability of the correct token | Error, approximately |
|---|---|
| 0.9 | 0.105 |
| 0.5 | 0.693 |
| 0.1 | 2.303 |
| 0.01 | 4.605 |
You can already see the effect in the table: a confident miss is penalized more heavily. As the correct probability grows, the error shrinks.
What does logarithm mean?
From softmax you know the exponential function exp. The natural logarithm is its inverse. Because exp(0)=1, ln(1)=0. Because exp(1)≈2.718, ln(2.718) is about 1.
For probabilities between 0 and 1, the natural logarithm is negative. The minus sign in front turns it into a positive error. In this book, log and ln in these formulas mean the same natural logarithm; in the Python code, the library computes it in a numerically stable way.
You don't have to work out the values in your head. What matters is understanding the direction and knowing a few values from the table. In the experiment you can follow the curve by changing the probability.
The short formula
Read it like this: pick out the probability of the correct token. Apply the natural logarithm to it and flip the sign. The result is the error L. k is the number of the correct token, and its probability. k is not a new learnable number.
For a correct probability of 0.5, the error is about 0.693. This rule for a single position is called the negative log-likelihood loss. For now, you can read that long name as “error from the probability of the correct token”.
Why it is also called cross-entropy
Cross-entropy is the general name for a related way of scoring distributions. In our case the target is a single correct token: conceptually, this class gets target weight one and all others zero. Such a representation is called a one-hot target. With it, cross-entropy simplifies to exactly the one term we have just computed.
In code we don't need to build this large one-hot vector. We simply pass the ID of the correct token. From that, PyTorch knows the right output position.
loss = F.cross_entropy(
logits.reshape(-1, vocab_size),
targets.reshape(-1),
ignore_index=-100,
)logits are the raw scores before softmax. targets contains the correct token IDs. Here, reshape groups batch and text positions into one shared list of prediction tasks. The vocabulary axis is kept. In this reshaping, -1 means: the library works out the matching length itself.
Internally, the function computes a stable combination of softmax and logarithm. So pass in the raw scores, not probabilities you have already converted. Otherwise you would be doing a different calculation.
Combining several positions fairly
We add up the errors of all active target positions and divide by their number. If one group has 100 targets and another has 10, the first group must get ten times the weight for a shared per-token average. Simply averaging the two group averages equally would be something different.
Padding means filler positions that give short examples the same length as the others in the batch. We don't want to learn these filler positions as if they were real text targets. So the code sets their targets to the special ignore value −100. ignore_index=-100 tells the loss function to leave these positions out. −100 is not a normal token ID.
An extra number called perplexity
Perplexity is what you get when you apply the exponential function to the mean error. With a mean error of ln(4), the perplexity is 4. It is a different scale for the same average, not a percentage of correct answers.
A uniform model with four possible tokens gives every correct token probability 1/4. Its mean error is −ln(1/4)=ln(4), so its perplexity is 4. This small uniform distribution is a useful reference point. You cannot use it to compare models with different tokenizers without further thought, because their units differ in size.
How does learning come out of this?
The loss function is connected to the raw scores and, through them, to the weights. Backpropagation can therefore compute its derivatives through the network, just as in the small number example. The exact derivative with respect to a score is “computed probability minus target value 0 or 1”. For the correct token, for example, this gives 0.5−1=−0.5, which locally pushes its score up.
You don't need to derive this particular derivative formula for your first training run. What matters is the connection: a lower error on the correct token can be reached through the same learning process you already know. The learning goal is new; the basic principle of training stays the same.