Tokenwerk · The LLM textbook

Chapter 10 · II · From numbers to language · 5 minutes

Your Python workshop

A small, understandable setup on Mac, Windows or Linux. First a reliable way to compute, then larger models.

Why Python and PyTorch?

The website is built with Vue. We program the actual model in Python, because PyTorch can compute with arrays of numbers, work out derivatives automatically and use compute chips such as graphics cards. PyTorch takes care of multiplications and differentiation (computing derivatives), but not our architecture, the tokenizer or the data decisions. You build the model classes yourself instead of loading a finished model with from_pretrained.

You don't need to have mastered Python. Lists, dictionaries, functions, classes and loops are enough to get started. Your own model class inherits from nn.Module. Weights you create as attributes (nn.Parameter) or as submodules (e.g. nn.Linear) are registered by the module automatically. model.parameters() then returns all registered parameters, which are normally the trainable numbers. A local variable that is only a tensor does not automatically become a parameter: that's what nn.Parameter is for.

Setting up the environment

Install Python 3.11 or 3.12. Use a virtual environment so that this project doesn't change the packages of other projects. Download the project package via the download in the textbook and unzip it.

bash
# macOS / Linux
cd tokenwerk-project
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt

In PowerShell, activation works differently:

powershell
py -3.12 -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -r requirements.txt

If PowerShell blocks activation, you can call the interpreter directly: .\.venv\Scripts\python.exe -m pip install -r requirements.txt. You don't need to change a system-wide policy just to run this project.

For an NVIDIA GPU you need a PyTorch build that matches your system; on Windows, pip install torch installs a CPU-only version. Use the official install selector for this. The simple installation in requirements.txt is no guarantee for every CUDA configuration. After installing, check:

python
import torch
print(torch.__version__)
print('CUDA:', torch.cuda.is_available())
print('MPS:', torch.backends.mps.is_available())
x = torch.tensor([1., 2., 3.], requires_grad=True)
(x.square().sum()).backward()
print(x.grad)  # tensor([2., 4., 6.])

Your Mac as a learning machine

On an Apple silicon Mac, PyTorch provides an accelerator through mps. For small tests the CPU is often the simpler and even faster starting point: on an Apple silicon Mac, 200 training steps of our small model took 1.3 seconds on the CPU and 4.1 seconds with mps – with identical loss. MPS only pays off with larger models and batches. Apple's own framework MLX is an alternative to PyTorch; this book and the project use PyTorch, so on a Mac that means cpu or mps. A Mac with 16 GB of unified memory is suitable for the learning path; but not all of that memory is available to the model. The operating system, programs, weights, gradients, optimizer and intermediate values all share it. Start with the micro configuration, not with billions of parameters.

python
def device_name():
    if torch.cuda.is_available():
        return 'cuda'
    if torch.backends.mps.is_available():
        return 'mps'
    return 'cpu'

An operation on an accelerator is not automatically faster. Small matrix multiplications can be limited by transfer and launch costs. A fair timing comparison needs warm-up and synchronization; otherwise a single unprepared timestamp may only measure the queuing of an operation.

Optional check of the complete project

bash
python tests.py
python prepare.py --demo --out runs/demo --vocab-size 320
python train.py --data runs/demo --out runs/micro --steps 100 --device cpu
python sample.py --checkpoint runs/micro/last.pt --prompt "The cat" --tokens 60

The technical terms that follow belong to later chapters. You don't need to understand or run the full pipeline yet; this is an optional installation test.

Among other things, tests.py checks tokenizer round trips, tensor shapes, causality, finite loss and the SFT mask. The 100 steps are a short functional test. If you expect good text afterwards, you are mistaking a running pipeline for a fully trained model.

For later: the project files

File Job
tokenizer.py Learn, save, load and apply byte-level BPE
model.py Classic and modern decoder
prepare.py Split documents and produce token sequences
train.py Pretraining, validation and checkpoints
sft.py Train on answer examples with masking
sample.py Generate text or an answer autoregressively
tests.py Small, targeted functional checks

A checkpoint contains more than weights. To resume, you also need the architecture configuration, the optimizer state, the step count and, ideally, the random states. The tokenizer is part of a model's identity: a different tokenizer with the same IDs can give the same weights completely different meanings. That's why the project also stores the tokenizer data in the checkpoint.

Small steps instead of long blind runs

For now, the installation and the derivative test are enough. In PyTorch an array of numbers is called a tensor; square() squares its values, sum() adds them up and backward() computes the derivatives. [1,2,3] becomes the squares [1,4,9], their sum 14 and the derivatives [2,4,6]. The next chapter explains such arrays and their shapes in detail.

Once you have worked through the later model chapters: run the tests first. Then check a small group of examples (a batch), one forward pass and one backward pass. After that, train on a few examples until the model can overfit them. Only when that works should you increase the amount of data. A bug in the target shift is not fixed by more compute. It just gets more expensive.