Tokenwerk · The LLM textbook

Chapter 23 · IV · Training and evaluating · 4 minutes

Reading the loss and finding bugs

A good-looking training curve can deceive you. With controlled tests you can detect leakage, underfitting and overfitting.

Four typical patterns

If training and validation loss both stay high, the model hasn't learned the task well yet. Possible reasons are too little training, an unsuitable learning rate, too little capacity or an implementation bug. If only training goes down and validation goes up, overfitting is plausible. If both are extremely good even without training, look for a target that is too easy or for leakage.

If both go down, that is an encouraging signal, but not final proof of answer quality. The token loss measures an average. A rare, important factual question can be answered wrongly even though many frequent function words are predicted correctly.

The diagnosis view in the browser shows schematic example curves, not results from a real training run. It is meant to help you choose suitable next checks. Your real measurements will later be in metrics.jsonl.

Overfitting one batch

Take a fixed small batch and train repeatedly on it alone. A sufficiently flexible model should be able to lower its loss clearly. If it can't, check the data path and the optimization first. For this test it is allowed to overfit; that is exactly the goal.

A successful overfitting test shows that the system can represent and optimize these examples. It does not show that it generalizes to new documents. So you shouldn't celebrate this step as a finished model and then call the same batch "validation".

Debugging in a sensible order

First check the raw text and the tokenizer round trip. Then check the input and the target shifted by one. Then check tensor shapes, the number of active targets and the causal prefix test. Only when these contracts hold should you examine the optimizer, learning rate and capacity.

This order keeps you from trying to compensate for a data bug with ever larger networks. A target shifted by two positions can make training harder; an unshifted target can make it seem easier. Both cases need the same early check.

Non-finite values

A NaN loss is not a normal learning state. Look for the first non-finite tensor. Are the inputs valid IDs? Do some attention rows contain only forbidden entries? Does a division by zero or an unstable exponential cause problems? Is the learning rate so large that weights explode?

python
if not torch.isfinite(loss):
    raise RuntimeError('Non-finite loss; run stopped.')
for name, p in model.named_parameters():
    if p.grad is not None and not torch.isfinite(p.grad).all():
        print('Non-finite gradient:', name)

Continuing to train unchecked can destroy the last usable checkpoint. Stop with a clear message and keep the previous checkpoint.

Training improves, sampling stays bad

First check the prompt in the actual tokenizer and the output format. A base model expects a text continuation, not an automatically recognized chat role. Also check the temperature and the token budget. A temperature that is too high can draw implausible tokens more often.

Compare generations with a fixed seed and several fixed prompts. One especially nice example out of a hundred attempts is not a reliable quality measurement. If you swap out the worst prompt after every model change, you are moving the test.

The train/validation gap

A gap is normal: training examples were optimized, validation examples weren't. The size and development of the gap are more informative than its mere existence. A larger gap can come from repeated data, distribution differences or different context handling. "That's overfitting" is therefore a hypothesis that you back up with data checks and controlled experiments.

Seeds and repetitions

A fixed seed helps make changes more comparable. It does not guarantee identical results on all devices. GPU kernels, parallel reductions and different library versions can produce differences. When results between two architectures are close, you should look at several seeds instead of interpreting a random ordering as a law of nature.

Your measurement sheet

Record the checkpoint, tokenizer ID or hash, data manifest, configuration, number of processed tokens, train loss, validation method, validation loss, runtime and fixed generation samples. For answer models, add criteria for factual quality.

A simple JSONL log is enough to start with. You don't need to integrate a big tracking service before your first model predicts a valid sequence. That's why the project writes its measurements locally to a simple file.

Next steps for each finding

Finding First sensible check
Loss unchanged Are weights learning? Are there gradients?
Near-perfect loss from the start Input/target identical? Future visible?
NaN after a few steps Learning rate, norms, first non-finite operation
Only training improves Data split, amount of data, early stopping
Good losses, bad answers Target distribution, chat format, separate evaluation

Change only one cause per experiment where possible. Several simultaneous interventions can improve a run, but afterwards you don't know which one helped.