Why Python and PyTorch?
The website is built with Vue. We program the actual model in Python, because PyTorch can compute with arrays of numbers, work out derivatives automatically and use compute chips such as graphics cards. PyTorch takes care of multiplications and differentiation (computing derivatives), but not our architecture, the tokenizer or the data decisions. You build the model classes yourself instead of loading a finished model with from_pretrained.
You don't need to have mastered Python. Lists, dictionaries, functions, classes and loops are enough to get started. Your own model class inherits from nn.Module. Weights you create as attributes (nn.Parameter) or as submodules (e.g. nn.Linear) are registered by the module automatically. model.parameters() then returns all registered parameters, which are normally the trainable numbers. A local variable that is only a tensor does not automatically become a parameter: that's what nn.Parameter is for.
Setting up the environment
Install Python 3.11 or 3.12. Use a virtual environment so that this project doesn't change the packages of other projects. Download the project package via the download in the textbook and unzip it.
# macOS / Linux
cd tokenwerk-project
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txtIn PowerShell, activation works differently:
py -3.12 -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -r requirements.txtIf PowerShell blocks activation, you can call the interpreter directly: .\.venv\Scripts\python.exe -m pip install -r requirements.txt. You don't need to change a system-wide policy just to run this project.
For an NVIDIA GPU you need a PyTorch build that matches your system; on Windows, pip install torch installs a CPU-only version. Use the official install selector for this. The simple installation in requirements.txt is no guarantee for every CUDA configuration. After installing, check:
import torch
print(torch.__version__)
print('CUDA:', torch.cuda.is_available())
print('MPS:', torch.backends.mps.is_available())
x = torch.tensor([1., 2., 3.], requires_grad=True)
(x.square().sum()).backward()
print(x.grad) # tensor([2., 4., 6.])Your Mac as a learning machine
On an Apple silicon Mac, PyTorch provides an accelerator through mps. For small tests the CPU is often the simpler and even faster starting point: on an Apple silicon Mac, 200 training steps of our small model took 1.3 seconds on the CPU and 4.1 seconds with mps – with identical loss. MPS only pays off with larger models and batches. Apple's own framework MLX is an alternative to PyTorch; this book and the project use PyTorch, so on a Mac that means cpu or mps. A Mac with 16 GB of unified memory is suitable for the learning path; but not all of that memory is available to the model. The operating system, programs, weights, gradients, optimizer and intermediate values all share it. Start with the micro configuration, not with billions of parameters.
def device_name():
if torch.cuda.is_available():
return 'cuda'
if torch.backends.mps.is_available():
return 'mps'
return 'cpu'An operation on an accelerator is not automatically faster. Small matrix multiplications can be limited by transfer and launch costs. A fair timing comparison needs warm-up and synchronization; otherwise a single unprepared timestamp may only measure the queuing of an operation.
Optional check of the complete project
python tests.py
python prepare.py --demo --out runs/demo --vocab-size 320
python train.py --data runs/demo --out runs/micro --steps 100 --device cpu
python sample.py --checkpoint runs/micro/last.pt --prompt "The cat" --tokens 60The technical terms that follow belong to later chapters. You don't need to understand or run the full pipeline yet; this is an optional installation test.
Among other things, tests.py checks tokenizer round trips, tensor shapes, causality, finite loss and the SFT mask. The 100 steps are a short functional test. If you expect good text afterwards, you are mistaking a running pipeline for a fully trained model.
For later: the project files
| File | Job |
|---|---|
tokenizer.py |
Learn, save, load and apply byte-level BPE |
model.py |
Classic and modern decoder |
prepare.py |
Split documents and produce token sequences |
train.py |
Pretraining, validation and checkpoints |
sft.py |
Train on answer examples with masking |
sample.py |
Generate text or an answer autoregressively |
tests.py |
Small, targeted functional checks |
A checkpoint contains more than weights. To resume, you also need the architecture configuration, the optimizer state, the step count and, ideally, the random states. The tokenizer is part of a model's identity: a different tokenizer with the same IDs can give the same weights completely different meanings. That's why the project also stores the tokenizer data in the checkpoint.
Small steps instead of long blind runs
For now, the installation and the derivative test are enough. In PyTorch an array of numbers is called a tensor; square() squares its values, sum() adds them up and backward() computes the derivatives. [1,2,3] becomes the squares [1,4,9], their sum 14 and the derivatives [2,4,6]. The next chapter explains such arrays and their shapes in detail.
Once you have worked through the later model chapters: run the tests first. Then check a small group of examples (a batch), one forward pass and one backward pass. After that, train on a few examples until the model can overfit them. Only when that works should you increase the amount of data. A bug in the target shift is not fixed by more compute. It just gets more expensive.