First we count four characters
A bigram is a pair of tokens that follow each other directly. A bigram model uses only the immediately preceding token to predict the next one. For this first example, every character is a token.
In the text abac we see the pairs a→b, b→a and a→c. After a, b came once and c came once. So we estimate: after a, b and c each get probability 1/2. The very last position has no next token here and therefore gives no additional pair.
| Previous | Next a | Next b | Next c | Row sum |
|---|---|---|---|---|
| a | 0 | 1 | 1 | 2 |
| b | 1 | 0 | 0 | 1 |
| c | 0 | 0 | 0 | 0 |
We call this count table C. The vocabulary size V is 3 here. So C has the shape [3,3], in general [V,V]. The first axis selects the previous token, the second the next one.
From a count to a probability
The calculation is: the number of times the desired transition occurred, divided by the number of all observed transitions after this previous token. In mathematical form:
| Symbol | Meaning |
|---|---|
| i | ID of the previous token, i.e. the row |
| j | ID of the desired next token, i.e. the column |
| Count in the cell at row i and column j | |
| k | Running index over all possible next tokens |
| Sum of the counts in row i | |
| Probability of j when i comes before it |
The notation describes the same idea with positions: xₜ is the token at position t, xₜ₋₁ the one immediately before it. The vertical bar means “given”. It has nothing to do with division.
Why we sometimes add artificial counts
For the c row we would be dividing by zero. Also, a continuation that was never observed should not necessarily be impossible forever. A simple fix: add the same small positive number to every cell, for example 1. This is called additive smoothing.
The a row [0,1,1] becomes [1,2,2]. The new sum is 5. The probabilities are [0.2,0.4,0.4]. The c row becomes [1,1,1] and gives 1/3 three times. Observed transitions are still preferred; unobserved ones are no longer ruled out.
All symbols mean the same as above. New are alpha (), the count added to each cell, and V, the number of possible next tokens. That is why alpha is added V times in the denominator. means “alpha is greater than zero”.
Smoothing is an assumption
With we pretend that every possible transition had one extra count. The larger alpha, the closer a sparsely observed row moves towards the uniform distribution. This protects against zero probabilities, but it can water down well-observed patterns. Alpha is not a “knowledge bonus”; it is a form of regularization for this estimate.
import torch
ids = torch.tensor([0, 1, 2, 2, 3])
V = 4
counts = torch.zeros(V, V)
for a, b in zip(ids[:-1], ids[1:]):
counts[a, b] += 1
prob = (counts + 1) / (counts + 1).sum(-1, keepdim=True)The browser lab uses real transition counts from your text. It is not a pretrained model and does not produce AI answers through a service.
Generating as a Markov chain
Choose a start token. Draw a next token from its row. Use it as the new previous token and repeat. This procedure often produces fragments that look right locally, but no reliable sentence structure. With a character model, “hello” can appear because the local transitions are plausible. The model can still produce “hellellell...”.
Its limitation is precise: two texts with the same last token lead to the same next distribution. “The capital of France is” and “The capital of Italy is” may well end with the same last token ID. A bigram model therefore cannot take the earlier country name into account. It does have context, but only one token of context.
A preview of the neural model
The count table needs no neural training. We count data and divide rows by their sum. Later we can use trainable numbers instead of the counts and compute their probabilities with softmax. Then an optimizer changes these numbers based on the error.
For that we are still missing two precise tools: embeddings as tables of numbers you can look up, and the right error measure for the correct next token ID. The next two chapters introduce them. Here it is enough to understand the counting bigram completely.
Why this baseline matters
Before you evaluate a transformer, you want to know whether your more complicated architecture is better than a simple model at all. Measure the bigram loss on the same validation documents with the same tokenizer. If your transformer is worse, optimization, data preparation or implementation may be wrong. It may also be that your dataset is too small or too simple to show a difference.
A baseline is not an opponent you should make artificially weak. Use a plausible smoothing and document it. The comparison is meant to make a decision easier, not to dramatize a result.
What a neural model adds
A trainable bigram table would have parameters: V rows times V columns. Our pure count table stores observed frequencies instead. For large vocabularies this grows quickly. An embedding with a small width plus a projection can share weights between contexts and later process several positions. The next learning step is not simply “more parameters”, but a structure that can sensibly share information between examples.
A controlled test
Train on a corpus with a completely unambiguous pattern, such as “abc” repeated. The model should learn the transitions a→b, b→c and c→a. Then enter “abd”. If d was never observed, smoothing or an untrained row takes over. From this small test you learn a lot about the input range, unknown symbols and your generation logic.