Time to build
So far you have met every component of a language model: tokens, embeddings, attention, blocks, loss, training, sampling. But you have only looked at them. In this part you write them yourself, in real PyTorch, on your own computer.
At the end there is a language model you wrote from the first line to the last. It learns from thousands of short children's stories and then writes stories of its own. After about two minutes of training on an ordinary laptop, it looks like this:
Once upon a time, there was a little boy named Timmy. Timmy liked to play
with his toys and play with him. One day, Timmy's mommy missed his mum and
daddy said, "Let's go play with Max!"Not exactly literature. But the model learned all by itself how these stories sound: “Once upon a time”, names, dialogue in quotation marks, moms and friends and toys. Nobody told it a single rule. And afterwards you even teach it to respond to requests.
The blueprint
We build in eight steps. Almost every step is one file, and every file turns ideas from a chapter you already know into code.
| Step | File | What it does | The idea is in |
|---|---|---|---|
| 1 | tokenizer.py |
Split text into token IDs and back | chapter 12 |
| 2 | prepare.py |
Load the stories, split them, turn them into IDs | chapter 21 |
| 3 | bigram.py |
A first, tiny neural language model | chapter 14 and 16 |
| 4 | model.py, check.py |
The GPT: attention, blocks, everything together | chapter 17 to 19 |
| 5 | train.py |
The training loop | chapter 22 |
| 6 | generate.py |
Generating text | chapter 24 |
| 7 | sft.py, chat.py |
The storyteller becomes a chat | chapter 28 |
| 8 | model.py |
Modernizing with RMSNorm, RoPE and SwiGLU | chapter 25 |
Together that is a little over 500 lines of Python, comments included. Sounds like a lot, but it is manageable: most files have 30 to 90 lines, and every piece is explained.
What you need
- A computer with Windows, macOS or Linux. You don't need a graphics card; everything runs on the normal processor.
- Python 3.10 or newer and about 1 GB of free space for PyTorch.
- Basic Python: variables, lists, loops, functions, classes. If you have written a small Python program before, that's enough.
- About an afternoon for all eight steps.
You don't have to type everything. The complete reference solution is available as a download: my-llm.zip. Honest tip: type the code yourself anyway. While typing, questions come up that you would never have while copying, and those make the difference.
Setting up the workshop
Create a new folder and set up a virtual environment in it (the details, including for Windows, are in chapter 10):
mkdir my-llm
cd my-llm
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install torch numpyThen a quick test that PyTorch runs:
python -c "import torch; print(torch.__version__)"If a version number like 2.6.0 appears, you are ready. Put all the files from the next chapters directly into this my-llm folder.
How you work in this part
Every build chapter follows the same pattern:
- What we build, and which concept chapter is behind it.
- The code, piece by piece, with every piece explained.
- The whole file for comparison.
- Run and check, with the output you should roughly see.
- If something goes wrong: the most common mistakes.
Your numbers won't look exactly like the ones in the book. Different computers, different PyTorch versions and randomness cause small differences. The order of magnitude should match, though. If something is way off, check the “If something goes wrong” section.
The training data: TinyStories
A language model needs text to learn from. We use TinyStories, a collection of short, simple children's stories that researchers created specifically to study small language models. The program downloads it by itself in the next step.
Why these stories? Three reasons: they use simple, consistent language, so a small model finds patterns quickly. They are clean: no web junk, no ads. And we use the validation part of the dataset, which is small enough to train on a laptop in minutes instead of days: about 19 MB of text, roughly 5.5 million tokens after splitting.
One thing up front, so you aren't surprised: the stories were generated by larger AI models, not written by people. That is exactly why their language is so uniform, and it makes them ideal for learning.
Key points
- In this part you write a complete GPT in PyTorch, in eight steps, mostly one file each.
- You only need Python, PyTorch and an ordinary computer, no graphics card.
- We train on TinyStories. In the end your model writes its own stories and responds to requests.