Tokenwerk · The LLM textbook

Chapter 34 · VII · Build your LLM · 8 minutes

Step 2: Preparing the data

Download TinyStories, split it fairly and turn it into long chains of numbers, all in one file.

What we build

Now we feed the tokenizer real text. The file prepare.py does everything chapter 21 recommended, in one go:

  1. Download TinyStories.
  2. Split the text into stories; every story is a document.
  3. First split the documents into training (90 %) and validation (10 %).
  4. Then learn the tokenizer only on the training documents.
  5. Turn both parts into token IDs, every document with <bos> in front and <eos> at the end.
  6. Save everything as files, so training can start right away later.

The order of points 3 and 4 is no accident. If the tokenizer also saw the validation texts, it would have a small unfair advantage when we check later: data leakage, chapter 21.

Loading the text

python
# Preparing the data: load text, split it, learn the tokenizer, turn everything into IDs (chapter 21)
import json
import random
import urllib.request
from pathlib import Path

import numpy as np

from tokenizer import Tokenizer

URL = "https://huggingface.co/datasets/roneneldan/TinyStories/resolve/main/TinyStories-valid.txt"
DATA = Path("data")
VOCAB_SIZE = 1024

The settings are at the top: where the text comes from, where the data goes and how big the vocabulary will be. Path("data") is simply the data folder next to your files.

python
def load_documents():
    raw = DATA / "tinystories.txt"
    if not raw.exists():
        DATA.mkdir(exist_ok=True)
        print("Downloading TinyStories (about 19 MB) …")
        urllib.request.urlretrieve(URL, raw)
    text = raw.read_text(encoding="utf-8")
    stories = [" ".join(s.split()) for s in text.split("<|endoftext|>")]   # one story per document
    return [s for s in stories if len(s) > 80]

What happens here:

  • On the first run urlretrieve downloads the text and saves it. After that it is stored locally, and the program never downloads it again.
  • In the raw file, stories are separated by a marker, <|endoftext|>. We split at it, so every story becomes one document.
  • Inside a story we turn all line breaks and multiple spaces into single spaces.
  • Stories under 80 characters are usually fragments. We drop them.

This is what the raw file looks like at a story boundary:

text
... They played together all day and became best friends.
<|endoftext|>
Once upon a time, in a big forest, there lived a rhinoceros named Roxy.
Roxy loved to climb. She climbed trees, rocks, and hills. ...

A clear separator like this is a gift: with many real texts you first have to figure out where one document ends and the next begins. Data quality is often manual work.

Splitting, learning the tokenizer, translating

python
def main():
    docs = load_documents()
    random.seed(42)
    random.shuffle(docs)
    n_valid = len(docs) // 10
    valid, train = docs[:n_valid], docs[n_valid:]                # split whole documents first ...
    print(f"{len(train)} training and {len(valid)} validation stories")
    (DATA / "train_docs.json").write_text(json.dumps(train, ensure_ascii=False), encoding="utf-8")

    tok = Tokenizer.train("\n".join(train), VOCAB_SIZE)          # ... then learn the tokenizer ONLY on training
    tok.save(DATA / "tokenizer.json")
    print("Vocabulary:", tok.vocab_size, "tokens")

    bos, eos = tok.special["<bos>"], tok.special["<eos>"]
    for name, split in [("train", train), ("valid", valid)]:
        ids = []
        for doc in split:
            ids += [bos] + tok.encode(doc) + [eos]               # every document with a start and an end
        np.save(DATA / f"{name}.npy", np.array(ids, dtype=np.int32))
        print(f"{name}: {len(ids):,} tokens")

    sample = tok.encode(train[0][:60])
    print("Example:", [tok.decode([i]) for i in sample])
  • random.seed(42) makes the “random” split the same on every run (chapter 21, seed).
  • train_docs.json remembers the training stories. We need them once more in step 7 for the chat.
  • [bos] + tok.encode(doc) + [eos] frames every document. That way the model later learns where a story starts and when to stop.
  • np.save stores the long ID lists as NumPy files. int32 is plenty, since our largest ID is 1023.

The whole file

python
# Preparing the data: load text, split it, learn the tokenizer, turn everything into IDs (chapter 21)
import json
import random
import urllib.request
from pathlib import Path

import numpy as np

from tokenizer import Tokenizer

URL = "https://huggingface.co/datasets/roneneldan/TinyStories/resolve/main/TinyStories-valid.txt"
DATA = Path("data")
VOCAB_SIZE = 1024


def load_documents():
    raw = DATA / "tinystories.txt"
    if not raw.exists():
        DATA.mkdir(exist_ok=True)
        print("Downloading TinyStories (about 19 MB) …")
        urllib.request.urlretrieve(URL, raw)
    text = raw.read_text(encoding="utf-8")
    stories = [" ".join(s.split()) for s in text.split("<|endoftext|>")]   # one story per document
    return [s for s in stories if len(s) > 80]


def main():
    docs = load_documents()
    random.seed(42)
    random.shuffle(docs)
    n_valid = len(docs) // 10
    valid, train = docs[:n_valid], docs[n_valid:]                # split whole documents first ...
    print(f"{len(train)} training and {len(valid)} validation stories")
    (DATA / "train_docs.json").write_text(json.dumps(train, ensure_ascii=False), encoding="utf-8")

    tok = Tokenizer.train("\n".join(train), VOCAB_SIZE)          # ... then learn the tokenizer ONLY on training
    tok.save(DATA / "tokenizer.json")
    print("Vocabulary:", tok.vocab_size, "tokens")

    bos, eos = tok.special["<bos>"], tok.special["<eos>"]
    for name, split in [("train", train), ("valid", valid)]:
        ids = []
        for doc in split:
            ids += [bos] + tok.encode(doc) + [eos]               # every document with a start and an end
        np.save(DATA / f"{name}.npy", np.array(ids, dtype=np.int32))
        print(f"{name}: {len(ids):,} tokens")

    sample = tok.encode(train[0][:60])
    print("Example:", [tok.decode([i]) for i in sample])


if __name__ == "__main__":
    main()

Run and check

bash
python prepare.py

The download (about 19 MB) takes a few seconds the first time, and the tokenizer about 15 seconds. You should see roughly this:

text
Downloading TinyStories (about 19 MB) …
19791 training and 2198 validation stories
Vocabulary: 1024 tokens
train: 5,515,701 tokens
valid: 607,454 tokens
Example: ['B', 'ell', 'a', ' was', ' th', 'ree', ' y', 'ear', 's', ' old', ...]

Look closely at the example: frequent words like “ was” or “ old” are a single token, rarer ones like “Bella” get split. That is exactly how BPE should work.

Now a look at the saved data. Start python and type:

python
import numpy as np
from tokenizer import Tokenizer
tok = Tokenizer.load("data/tokenizer.json")
train = np.load("data/train.npy")
print(len(train))
print(train[:10])
print(tok.decode(train[:30]))
text
5515701
[1020   66  574   97  280  304  898  332  822  115]
Bella was three years old and excited to go to the park. She had been watching more birds and

The first number, 1020, is <bos>: 256 bytes plus 763 merges give 1019 normal tokens (IDs 0 to 1018), followed by <pad> (1019) and <bos> (1020). The rest are the words of the first story. This is how your model sees the world: one long chain of numbers.

If something goes wrong

  • URLError or a timeout during download: no internet connection, or Hugging Face is briefly unreachable. Alternatively, download the file TinyStories-valid.txt in your browser from the dataset page and save it as data/tinystories.txt.
  • MemoryError or very slow: on a computer with very little memory, use only part of the stories, for example docs = load_documents()[:5000] in main.
  • Clearly different numbers: slightly different values are normal if the dataset was updated. Thousands of stories and several million tokens should be there, though.

Key points

  • prepare.py loads the text, splits it into stories and divides first into training and validation.
  • The tokenizer only learns on the training data. Every document is framed with <bos> and <eos>.
  • The result is two long chains of numbers in data/: about 5.5 million tokens for learning and 600,000 for checking.