A base model continues text
A base model is trained to continue text from its training distribution. Given "Why is ice cold?", it might produce another question, a dialogue marker or an explanation. That we expect a helpful answer is an extra requirement on its behavior. Supervised fine-tuning, or SFT for short, trains on examples with inputs and desired answers.
SFT can teach format, tone, brevity and task behavior. But a small model with little pretraining does not become a general knowledge system from a hundred question-answer pairs. It may memorize answers instead of transferring their content. That is why, in this course, we test narrowly scoped tasks and new phrasings of questions.
The chat template
Our project uses reserved IDs. Schematically, a single conversation looks like this:
BOS USER Frage EOS ASSISTANT Antwort EOSAn optional system text reads SYSTEM Text EOS after BOS. Further messages can continue the same structure. Roles and endings are inserted explicitly during data preparation, not guessed from arbitrary text fragments.
When generating an answer, the prompt ends after ASSISTANT. The next token should be the first answer token. If you use different templates for training and inference, you create a distribution shift that can seriously disrupt small models in particular.
Loss only on the answer
For this course project we ignore system and user targets as well as role markers. Answer tokens and the final EOS are active targets. Attention can still read context tokens; they just don't produce any loss at their own target positions.
The shift is what matters: if mask[j] says whether the token at position j is a target to learn, then with the input ids[:-1]:
targets = ids[1:].copy()
active = mask[1:]
targets[~active] = -100So the mask is shifted by one position, just like the target. The output at the ASSISTANT marker predicts the first answer token and therefore needs an active target.
A concrete format
{"messages":[{"role":"user","content":"Why does ice melt?"},{"role":"assistant","content":"Ice melts when enough heat is added to it."}]}The project includes small sft_train.jsonl and sft_valid.jsonl examples. They check that the pipeline works. For reliable answer quality you need far more varied, correct examples and a separate evaluation. The saved pretraining tokenizer stays unchanged.
python sft.py --checkpoint runs/base/best.pt \
--input data/sft_train.jsonl --valid data/sft_valid.jsonl \
--out runs/chat --steps 300 --lr 0.0001
python sample.py --checkpoint runs/chat/best.pt \
--chat --prompt "Why does ice melt?" --tokens 100The SFT trainer updates all model weights. It does not use LoRA adapters. For small models that is understandable and cheap enough; the point is to follow the entire learning path.
Don't blindly cut off long examples
If you keep only the first T+1 tokens, the entire answer can be lost. Then the example contains no active target. An averaged cross-entropy loss over nothing but ignored targets can become undefined.
Our project discards examples that don't fit in the context window instead of silently cutting off their answers. It reports how many were discarded and stops if no valid examples are left. In a large pipeline you could shorten messages intelligently or choose suitable answer segments; that has to fit the template and the learning goal.
Padding and token weighting
SFT examples vary in length. The batch is padded on the right with PAD, and the corresponding targets are -100. Because of causality, active positions never read later padding. The loss averages over the answer targets that are actually active.
During validation the project sums individual losses and counts active targets across all examples. Otherwise a short answer would get as much weight as a long one, even though we report a per-token average. A different weighting per example can be deliberately useful, but it is a different metric.
What the examples need to cover
If you want short English answers, the examples must show short English answers. Add different phrasings of questions, requests for explanations, simple text editing and situations without enough information. Check the desired answers yourself. With small amounts of data, contradictory patterns can have a strong effect.
Careful selection is no substitute for broad language data. You can tailor pretraining and SFT to a narrow topic; then describe the model as a narrow topic model too, and test how it responds outside that topic.
Compare before and after SFT
Use the same questions, fixed in advance, with the base and the SFT checkpoint. Rate format, clarity, correctness and appropriately admitting what it doesn't know. A better answer format with the same factual errors is real but limited progress. Record these two results separately.