The task
Build a small topic model in your own language or an English story model. You train the tokenizer and all decoder weights yourself. Then you add a small SFT run for short answers. The core of success is not a perfect sentence but a complete, tested path that you can explain.
The download package is a reference project. You may start with it, but you should be able to explain every central function and implement a small change yourself. "Fully on your own" here means: you can adapt data and architecture, narrow down errors and justify the next attempt, without relying on an opaque model-loading function.
Milestone A: the pipeline runs
python tests.py
python prepare.py --demo --out runs/demo --vocab-size 320
python train.py --data runs/demo --out runs/micro \
--device cpu --steps 100 --context 64 --dim 64 --layers 2 --heads 4
python sample.py --checkpoint runs/micro/last.pt \
--prompt "The cat" --tokens 60 --temperature 0.8Success: the tests pass, the trainer writes valid metrics and checkpoints, and the sampler decodes an output. This milestone does not yet require good language. The corpus is deliberately too small for that.
Milestone B: a real language-learning run
Create data/documents.jsonl from suitable, usable sources. Alternatively, you can start with the optional TinyStories helper script; it additionally needs the datasets package and network access. Check the dataset description and terms of use at the source.
python -m pip install datasets
python fetch_tinystories.py --count 5000 --out data/stories.jsonl
python prepare.py --input data/stories.jsonl --out runs/stories --vocab-size 1024
python train.py --data runs/stories --out runs/stories-base \
--modern --context 128 --dim 128 --layers 4 \
--heads 4 --kv-heads 2 --batch 8 --steps 2000This sample and these steps are a pilot, not a final configuration that is guaranteed to be sufficient. Evaluate loss and samples. For better text, you have to increase data, model and budget depending on what you find. TinyStories is in English, which fits English SFT. For another language, SFT on an English micro model alone is not a reliable path to ability in that language.
Milestone C: compare against the baseline
python baseline.py --data runs/stories
python evaluate.py --checkpoint runs/stories-base/best.pt --data runs/storiesBoth programs use the same tokenizer and the same prepared validation sequence. The baseline learns smoothed bigram transitions on the training data. The full decoder evaluation uses non-overlapping context blocks and reports a token-weighted loss; the context conditions necessarily differ from a bigram model that reads only one token. Document this method.
Success: with a suitable training run, the decoder should get below the bigram loss. If it doesn't, work through the diagnostics chapter. Also report ten fixed continuations and typical errors.
Milestone D: train answers
For a model in another language, you use pretraining and new SFT examples in that language. Split SFT training and SFT validation by questions or topics, not just by lines that practically rephrase the same question.
python sft.py --checkpoint runs/base/best.pt \
--input data/sft_train.jsonl --valid data/sft_valid.jsonl \
--out runs/chat --steps 300 --lr 0.0001
python sample.py --checkpoint runs/chat/best.pt \
--chat --prompt "Why does a plant need light?" \
--tokens 100 --temperature 0.6The included SFT files are small working examples. For a dependable model you replace or extend them. An overfitted demo run can answer known questions and fail on new ones. That is exactly why milestone E is part of the project.
Milestone E: independent evaluation
Create the 20 test questions and rate each answer with the rubric from the previous chapter. Compare base and SFT on the same questions and seeds. Note at least one success and three typical errors. Keep held-back questions away from training.
Success could mean, for example: on a narrow topic of your own, the model gives mostly understandable answers and improves its fit to the task through SFT. Set concrete thresholds that suit your goal. If factual correctness is still weak, that is a result, not a cue to make the test questions easier after the fact.
Milestone F: change something yourself
Pick exactly one extension. Implement a correct KV cache with a test that the logits are identical. Or compare two model widths with the same token budget. Or replace the BPE trainer with a more efficient implementation that passes identical round-trip tests. Document what your change measures and what it doesn't.
Your final report
Write two to four pages: goal and scope, data manifest, architecture, parameter count, training budget, evaluation methods, results, errors and the next sensible experiment. Put the tokenizer, the configuration and the fixed test list next to it. That way someone else can reproduce your result.
When you've reached the learning goal
You can explain how a character turns into a token and an embedding. You can compute an attention row. You can show why the mask is causal. You can justify the target shift and the SFT mask. You can start, stop, resume and evaluate a complete training run. And you can replace a claim about model quality with a suitable measurement.
That is the foundation on which you can keep building on your own. From here, frontier training is a large scaling and research task, but the computational mechanisms are no longer a black box to you.