What we build
Before we build the big GPT, there's a test run: the smallest neural language model there is. It is the bigram model from chapter 14, except that this time we don't count; we let the table learn with gradient descent.
Why the detour? Because along the way you meet the four PyTorch building blocks that are in every model, from a mini bigram to ChatGPT:
- Tensors: number fields with a shape (chapter 11).
nn.Module: a building block with learnable weights.F.cross_entropy: the error from chapter 16.- The optimizer: turns the weights downhill (chapter 8).
Once these work together, the rest is “just” a bigger model.
PyTorch in five minutes
A tensor is PyTorch's number field, much like a NumPy array. The big difference: PyTorch can remember how a result came about and compute the slopes backwards automatically. Try this in the Python console:
import torch
x = torch.tensor([[1.0, 2.0], [3.0, 4.0]])
print(x.shape)
print(x @ x)
w = torch.tensor(1.0, requires_grad=True)
loss = (w * 2 - 6) ** 2
loss.backward()
print(loss.item(), w.grad.item())torch.Size([2, 2])
tensor([[ 7., 10.],
[15., 22.]])
16.0 -16.0Recognize the last line? That is exactly the example from chapter 8: weight 1, input 2, target 6, error 16 and slope −16. This time PyTorch computed the derivative by itself: requires_grad=True says “remember the calculation path”, backward() computes backwards, and w.grad holds the slope. That is backpropagation from chapter 9, fully automatic.
The model
# The first trainable language model: a learnable bigram table (chapters 14–16)
import numpy as np
import torch
import torch.nn as nn
import torch.nn.functional as F
from tokenizer import Tokenizer
tok = Tokenizer.load("data/tokenizer.json")
train = torch.from_numpy(np.load("data/train.npy")).long()
V = tok.vocab_sizeWe load the tokenizer and the training data. torch.from_numpy(...).long() turns the NumPy array into a tensor of whole numbers; token IDs have to be whole numbers.
class Bigram(nn.Module):
def __init__(self):
super().__init__()
self.table = nn.Embedding(V, V) # row = previous token, column = score for the next one
def forward(self, ids):
return self.table(ids) # Logits: [B, T, V]That is the whole model. Every PyTorch class of your own inherits from nn.Module and has two parts:
- In
__init__you create the building blocks that have weights. Here just one:nn.Embedding(V, V), a table with 1024 rows and 1024 columns. Row = previous token, column = score for each possible next token. A little over a million learnable numbers in total. - In
forwardyou describe how the input turns into the output. Here: look up the matching row for every token ID. Those are already the logits.
Later you simply call the model with model(x); PyTorch then calls forward for you.
The training loop in miniature
torch.manual_seed(0)
model = Bigram()
optimizer = torch.optim.AdamW(model.parameters(), lr=0.1)
for step in range(301):
starts = torch.randint(0, len(train) - 33, (32,))
x = torch.stack([train[s : s + 32] for s in starts])
y = torch.stack([train[s + 1 : s + 33] for s in starts]) # Target-Shift
logits = model(x)
loss = F.cross_entropy(logits.view(-1, V), y.view(-1))
optimizer.zero_grad()
loss.backward()
optimizer.step()
if step % 50 == 0:
print(f"Step {step:3d} loss {loss.item():.3f}")This is the learning cycle from chapter 8, now with real text:
- Draw a batch: 32 random starting points, 32 tokens from each as input
x. The targetyis the same spot shifted by one: the target shift from chapter 16. Atx[i]the model should predicty[i], i.e. the next token each time. - Forward:
model(x)returns logits of shape[32, 32, 1024]: 32 texts, 32 positions, 1024 scores. - Error:
F.cross_entropywants two flat lists: all predictions stacked (view(-1, V)turns them into[1024, 1024]) and all targets (view(-1)). The function computes softmax and logarithm itself. - Backward and step:
zero_grad()clears the old slopes,backward()computes the new ones,step()turns all weights a bit downhill.
Generating text
ids = [tok.special["<bos>"]]
for _ in range(40):
probs = F.softmax(model(torch.tensor([ids]))[0, -1], dim=-1)
ids.append(torch.multinomial(probs, 1).item())
print(tok.decode(ids))We start with <bos> and ask 40 times: what are the probabilities of the next token? Then we roll one with torch.multinomial and append it. Just like the guessing game from chapter 1, except that the model learned the percentages itself.
The whole file
# The first trainable language model: a learnable bigram table (chapters 14–16)
import numpy as np
import torch
import torch.nn as nn
import torch.nn.functional as F
from tokenizer import Tokenizer
tok = Tokenizer.load("data/tokenizer.json")
train = torch.from_numpy(np.load("data/train.npy")).long()
V = tok.vocab_size
class Bigram(nn.Module):
def __init__(self):
super().__init__()
self.table = nn.Embedding(V, V) # row = previous token, column = score for the next one
def forward(self, ids):
return self.table(ids) # Logits: [B, T, V]
torch.manual_seed(0)
model = Bigram()
optimizer = torch.optim.AdamW(model.parameters(), lr=0.1)
for step in range(301):
starts = torch.randint(0, len(train) - 33, (32,))
x = torch.stack([train[s : s + 32] for s in starts])
y = torch.stack([train[s + 1 : s + 33] for s in starts]) # Target-Shift
logits = model(x)
loss = F.cross_entropy(logits.view(-1, V), y.view(-1))
optimizer.zero_grad()
loss.backward()
optimizer.step()
if step % 50 == 0:
print(f"Step {step:3d} loss {loss.item():.3f}")
ids = [tok.special["<bos>"]]
for _ in range(40):
probs = F.softmax(model(torch.tensor([ids]))[0, -1], dim=-1)
ids.append(torch.multinomial(probs, 1).item())
print(tok.decode(ids))Run and check
python bigram.pyStep 0 loss 7.474
Step 50 loss 4.513
Step 100 loss 3.998
Step 150 loss 3.814
Step 200 loss 3.720
Step 250 loss 3.728
Step 300 loss 3.763
Once upon a little follow. They see allnna and he knew it to really happy. They through he heard vasebow, but she heard some salse toHow to read this:
- It starts at about 7.5. A model guessing blindly would have ln(1024) ≈ 6.9. We start a bit worse because PyTorch fills the table with fairly large random numbers, so the model doesn't guess evenly but randomly wrong. We fix that in the GPT in a moment.
- After 300 steps we are at about 3.7. The model has learned which tokens typically follow each other. The text already looks like English: real words, punctuation, even “Once upon a”.
- But nothing makes sense. Of course: the model only ever sees the very last token. It has a memory of exactly one word, as described in chapter 14.
Remember the 3.7. That is your baseline: the GPT in the next step has to get clearly below it; otherwise something is wrong.
If something goes wrong
FileNotFoundError: data/tokenizer.json: runpython prepare.pyfirst.RuntimeError: shape '[-1, 1024]' is invalid: the vocabulary size doesn't match. Check that it saysV = tok.vocab_sizeand not a fixed number.- The loss doesn't go down: check the shift:
ymust start ats + 1, not ats. Andoptimizer.zero_grad()must come beforebackward().
Key points
- A PyTorch model is a class with
__init__(building blocks) andforward(calculation path). - Every training round: draw a batch, forward,
cross_entropy,zero_grad,backward,step. - The bigram model reaches a loss of about 3.7. That is the bar for the GPT.