One forward pass is not yet a paragraph
A decoder delivers logits for each input position. To generate, we use the logits of the last position, because they predict the next token after the entire visible prefix. We pick an ID, append it to the context and compute again.
with torch.inference_mode():
logits = model(ids[:, -cfg.context:])[:, -1, :]
next_id = torch.argmax(logits, dim=-1, keepdim=True)
ids = torch.cat([ids, next_id], dim=1)The weights don't change in the process. Without a KV cache, the visible context is processed again at every step. That is slow, but easy to check when you're starting out.
Greedy decoding
Greedy picks the token with the highest score. With the same model, prefix and numerical environment, this rule is largely deterministic. It does not automatically maximize the probability of the entire future sequence. A locally best continuation can later lead to an unfavorable sequence.
With small models, greedy can reinforce repetition loops. Randomness in sampling can sometimes break such loops, but it doesn't remove a structural weakness of the model. If a model rates "and and and" highly, you shouldn't just adjust its temperature.
Temperature and randomness
Divide the logits by a positive temperature and turn them into probabilities. Draw an ID with torch.multinomial. A random seed makes reproducible comparisons easier. A different seed doesn't produce different knowledge, only different draws from the computed distributions.
logits = logits / temperature
probs = torch.softmax(logits, dim=-1)
next_id = torch.multinomial(probs, num_samples=1)In the project, --temperature 0 means greedy; the code then doesn't divide by zero. Negative temperatures are rejected.
Top-k
Top-k allows only the k largest logits. All others get minus infinity. Then the distribution is normalized again. With k=1 this is the same as picking the largest score, apart from questions of ties. A small k can suppress implausible fringe tokens, but it also limits variety.
k = min(top_k, logits.size(-1))
threshold = torch.topk(logits, k).values[..., -1, None]
logits = logits.masked_fill(logits < threshold, float('-inf'))With exactly tied scores, this threshold form can keep more than k tokens. The reference project uses the explicit top-k indices to keep exactly the chosen number.
Top-p or nucleus sampling
Sort the tokens by probability. Keep the smallest leading set whose cumulative probability reaches at least p. The token that crosses the threshold is included. This way the number of allowed tokens adapts to how concentrated the distribution is.
For probabilities [0.6,0.25,0.1,0.05] and p=0.8, we keep the first two tokens. With p=0.95 it's three. In the implementation, watch out for rounding and make sure to keep at least one token. Applying top-k and top-p one after the other gives a different distribution than using only one of them; document the order.
EOS is a trained stop
In pretraining, the model sees EOS after documents. In SFT, it sees EOS after an answer. During generation, we stop when EOS is chosen or when a maximum token budget is reached. Without EOS training, outputs can result that never end sensibly.
We don't want to produce reserved role markers in the middle of the visible answer text. Our sampler blocks PAD, BOS and role markers in normal output but allows EOS. This is a decoding rule, not a claim that the model itself would never score these classes highly.
The context window and cut-off memory
If the context is longer than the configured length, our simple sampler takes only the last tokens. The model can then read only those tokens. An early system marker or a question can drop out. A long output therefore doesn't mean that all earlier inputs are still taken into account.
In the course project, prompts and the answer budget should stay short. A serious chat application needs a deliberate strategy for shortening messages, not just a blind window over the entire conversation.
How to check generation
python sample.py --checkpoint runs/base/best.pt \
--prompt "Ein kleiner Hund" --tokens 100 \
--temperature 0.8 --top-k 40 --top-p 0.95 --seed 42Compare the same prompts with greedy and two temperatures. Check readability, repetition, stopping and format. Don't publish only the nicest sample. With a demonstration corpus, memorization is to be expected; with a larger corpus, also examine similarity to training texts.
The experiment
Draw a token several times from a three-part distribution. After a few draws, the observed frequencies need not match the theoretical probabilities exactly. Over many draws they should approach them. This random fluctuation explains why individual samples are so unreliable for comparing models.