An architecture learns the distribution of your examples
A small model that only ever sees short children's stories will not automatically answer technical support questions. A model that only sees tables of questions and answers may learn the format but little general language. Your data largely define the range in which generalization is plausible at all.
For the first run we use the included demonstration corpus. It contains short English examples we wrote ourselves. It is small and repeats simple structures. It is meant for checking that everything works, not as a promise of performance. The download also includes a script to optionally fetch texts from TinyStories. TinyStories is in English and studies small language models on simplified stories; that makes it interesting for language training, but not a ready-made dialogue dataset.
Split documents first
Picture a novel that you cut into overlapping windows. If you distribute the windows randomly between train and validation, both sets contain almost identical sentences. The validation loss then looks better than a test on genuinely new documents would.
Our order is: read the raw documents, remove exact duplicates, split the documents into training and validation with a fixed seed, learn the tokenizer on training only, tokenize both sets. For a real release you also need an independent test set that was not used to choose hyperparameters.
A simple file format
prepare.py expects JSONL: one JSON line per document.
{"text":"Lina finds a small key. She looks for the matching door."}
{"text":"A thunderstorm forms when moist air rises strongly."}One field per document makes the boundaries explicit. A single gigantic text file without boundaries is harder to split cleanly. You can convert your own sources into this format, but you should document their origin and rights separately.
python prepare.py --input data/documents.jsonl --out runs/data --vocab-size 2048The naive teaching BPE training can become slow on large corpora. Our script limits the data on which merges are counted. This is a documented sample of the training data, not a use of the validation set. For millions of documents, you will later replace this part with an optimized tokenizer trainer, once you understand the learning logic.
What a corpus manifest should contain
For each source, note its origin, language, date, license or basis for use, number of documents, approximate token count, filter rules and known limitations. Whether reuse is actually permitted depends on the source and your use case; our course makes no blanket legal statement about other people's texts.
Watch out for personal data and secrets in your own datasets. A small model in particular can memorize heavily. Don't thoughtlessly train on internal credentials, support tickets or private messages. Remove such content using clear rules and check samples of the filtered result.
Quality is more than grammatically clean text
A corpus can have flawless grammar and still contain false facts. Or it contains a great many nearly identical SEO texts. A simple quality check looks for empty texts, unusual character proportions, very long repetitions, identical documents and rough language identification. Exact deduplication is not enough to detect slightly modified copies.
Near-duplicate detection and high-quality fact checking are large topics of their own. For your learning project, you document this limitation instead of calling rudimentary text cleaning "complete quality control".
EOS and sequence packing
Our pretraining appends EOS after every document and puts the tokens of a set into one shared sequence. Randomly drawn training windows can cross a document boundary. Causal attention may then also look at text before EOS. EOS is a learned boundary signal, not a technical reset of all hidden states.
This is a deliberately simple packing method. If documents should be fully isolated from each other, you need document-aware attention masks or separate sequences. For the course project, shared packing is manageable, but you should know what it means.
Inspect samples before training
Decode twenty randomly selected windows. Are special characters intact? Are there empty documents or accidentally merged JSON syntax? Are EOS boundaries visible? Then show the input and target of the same window side by side. Actually verify the shift before you start a long run.
Token budget instead of a "feel for epochs"
If you draw random windows with replacement, an "epoch" is not automatically one complete pass through every token. That's why our trainer records steps and processed training tokens. A small corpus can be reused many times. New steps then don't create new language examples; they keep optimizing on the same data.
The experiment
The split view shows different documents and how they are assigned with a fixed seed. A second view shows why overlapping windows can break the separation. The view does not produce a quality score for your real dataset. It helps you check the order of the processing steps.