What we build
Now we feed the tokenizer real text. The file prepare.py does everything chapter 21 recommended, in one go:
- Download TinyStories.
- Split the text into stories; every story is a document.
- First split the documents into training (90 %) and validation (10 %).
- Then learn the tokenizer only on the training documents.
- Turn both parts into token IDs, every document with
<bos>in front and<eos>at the end. - Save everything as files, so training can start right away later.
The order of points 3 and 4 is no accident. If the tokenizer also saw the validation texts, it would have a small unfair advantage when we check later: data leakage, chapter 21.
Loading the text
# Preparing the data: load text, split it, learn the tokenizer, turn everything into IDs (chapter 21)
import json
import random
import urllib.request
from pathlib import Path
import numpy as np
from tokenizer import Tokenizer
URL = "https://huggingface.co/datasets/roneneldan/TinyStories/resolve/main/TinyStories-valid.txt"
DATA = Path("data")
VOCAB_SIZE = 1024The settings are at the top: where the text comes from, where the data goes and how big the vocabulary will be. Path("data") is simply the data folder next to your files.
def load_documents():
raw = DATA / "tinystories.txt"
if not raw.exists():
DATA.mkdir(exist_ok=True)
print("Downloading TinyStories (about 19 MB) …")
urllib.request.urlretrieve(URL, raw)
text = raw.read_text(encoding="utf-8")
stories = [" ".join(s.split()) for s in text.split("<|endoftext|>")] # one story per document
return [s for s in stories if len(s) > 80]What happens here:
- On the first run
urlretrievedownloads the text and saves it. After that it is stored locally, and the program never downloads it again. - In the raw file, stories are separated by a marker,
<|endoftext|>. We split at it, so every story becomes one document. - Inside a story we turn all line breaks and multiple spaces into single spaces.
- Stories under 80 characters are usually fragments. We drop them.
This is what the raw file looks like at a story boundary:
... They played together all day and became best friends.
<|endoftext|>
Once upon a time, in a big forest, there lived a rhinoceros named Roxy.
Roxy loved to climb. She climbed trees, rocks, and hills. ...A clear separator like this is a gift: with many real texts you first have to figure out where one document ends and the next begins. Data quality is often manual work.
Splitting, learning the tokenizer, translating
def main():
docs = load_documents()
random.seed(42)
random.shuffle(docs)
n_valid = len(docs) // 10
valid, train = docs[:n_valid], docs[n_valid:] # split whole documents first ...
print(f"{len(train)} training and {len(valid)} validation stories")
(DATA / "train_docs.json").write_text(json.dumps(train, ensure_ascii=False), encoding="utf-8")
tok = Tokenizer.train("\n".join(train), VOCAB_SIZE) # ... then learn the tokenizer ONLY on training
tok.save(DATA / "tokenizer.json")
print("Vocabulary:", tok.vocab_size, "tokens")
bos, eos = tok.special["<bos>"], tok.special["<eos>"]
for name, split in [("train", train), ("valid", valid)]:
ids = []
for doc in split:
ids += [bos] + tok.encode(doc) + [eos] # every document with a start and an end
np.save(DATA / f"{name}.npy", np.array(ids, dtype=np.int32))
print(f"{name}: {len(ids):,} tokens")
sample = tok.encode(train[0][:60])
print("Example:", [tok.decode([i]) for i in sample])random.seed(42)makes the “random” split the same on every run (chapter 21, seed).train_docs.jsonremembers the training stories. We need them once more in step 7 for the chat.[bos] + tok.encode(doc) + [eos]frames every document. That way the model later learns where a story starts and when to stop.np.savestores the long ID lists as NumPy files.int32is plenty, since our largest ID is 1023.
The whole file
# Preparing the data: load text, split it, learn the tokenizer, turn everything into IDs (chapter 21)
import json
import random
import urllib.request
from pathlib import Path
import numpy as np
from tokenizer import Tokenizer
URL = "https://huggingface.co/datasets/roneneldan/TinyStories/resolve/main/TinyStories-valid.txt"
DATA = Path("data")
VOCAB_SIZE = 1024
def load_documents():
raw = DATA / "tinystories.txt"
if not raw.exists():
DATA.mkdir(exist_ok=True)
print("Downloading TinyStories (about 19 MB) …")
urllib.request.urlretrieve(URL, raw)
text = raw.read_text(encoding="utf-8")
stories = [" ".join(s.split()) for s in text.split("<|endoftext|>")] # one story per document
return [s for s in stories if len(s) > 80]
def main():
docs = load_documents()
random.seed(42)
random.shuffle(docs)
n_valid = len(docs) // 10
valid, train = docs[:n_valid], docs[n_valid:] # split whole documents first ...
print(f"{len(train)} training and {len(valid)} validation stories")
(DATA / "train_docs.json").write_text(json.dumps(train, ensure_ascii=False), encoding="utf-8")
tok = Tokenizer.train("\n".join(train), VOCAB_SIZE) # ... then learn the tokenizer ONLY on training
tok.save(DATA / "tokenizer.json")
print("Vocabulary:", tok.vocab_size, "tokens")
bos, eos = tok.special["<bos>"], tok.special["<eos>"]
for name, split in [("train", train), ("valid", valid)]:
ids = []
for doc in split:
ids += [bos] + tok.encode(doc) + [eos] # every document with a start and an end
np.save(DATA / f"{name}.npy", np.array(ids, dtype=np.int32))
print(f"{name}: {len(ids):,} tokens")
sample = tok.encode(train[0][:60])
print("Example:", [tok.decode([i]) for i in sample])
if __name__ == "__main__":
main()Run and check
python prepare.pyThe download (about 19 MB) takes a few seconds the first time, and the tokenizer about 15 seconds. You should see roughly this:
Downloading TinyStories (about 19 MB) …
19791 training and 2198 validation stories
Vocabulary: 1024 tokens
train: 5,515,701 tokens
valid: 607,454 tokens
Example: ['B', 'ell', 'a', ' was', ' th', 'ree', ' y', 'ear', 's', ' old', ...]Look closely at the example: frequent words like “ was” or “ old” are a single token, rarer ones like “Bella” get split. That is exactly how BPE should work.
Now a look at the saved data. Start python and type:
import numpy as np
from tokenizer import Tokenizer
tok = Tokenizer.load("data/tokenizer.json")
train = np.load("data/train.npy")
print(len(train))
print(train[:10])
print(tok.decode(train[:30]))5515701
[1020 66 574 97 280 304 898 332 822 115]
Bella was three years old and excited to go to the park. She had been watching more birds andThe first number, 1020, is <bos>: 256 bytes plus 763 merges give 1019 normal tokens (IDs 0 to 1018), followed by <pad> (1019) and <bos> (1020). The rest are the words of the first story. This is how your model sees the world: one long chain of numbers.
If something goes wrong
URLErroror a timeout during download: no internet connection, or Hugging Face is briefly unreachable. Alternatively, download the fileTinyStories-valid.txtin your browser from the dataset page and save it asdata/tinystories.txt.MemoryErroror very slow: on a computer with very little memory, use only part of the stories, for exampledocs = load_documents()[:5000]inmain.- Clearly different numbers: slightly different values are normal if the dataset was updated. Thousands of stories and several million tokens should be there, though.
Key points
prepare.pyloads the text, splits it into stories and divides first into training and validation.- The tokenizer only learns on the training data. Every document is framed with
<bos>and<eos>. - The result is two long chains of numbers in
data/: about 5.5 million tokens for learning and 600,000 for checking.