Tokenwerk · The LLM textbook

Chapter 12 · II · From numbers to language · 5 minutes

From text to token IDs

Characters, bytes and subwords: how to turn text into indices without losing anything, and how to train your own BPE tokenizer.

IDs are addresses, not quantities

A language model processes numbers. The ID 42 does not mean “more” than the ID 17. It is an address in a table. You could rename all IDs with a permutation and reorder the tables to match; the model would compute the same thing. That is why you must not simply feed the IDs into a linear layer as scalar numbers.

A token can be a character, a byte, a piece of a word, a whole word or a control symbol. Tokens are not automatically words. Tokenization decides how the text is split into the positions at which the model will later make its predictions.

Character tokenizer: quick to understand

python
text = 'Hallo Welt!'
vocab = sorted(set(text))
stoi = {c: i for i, c in enumerate(vocab)}
itos = {i: c for c, i in stoi.items()}
ids = [stoi[c] for c in text]
assert ''.join(itos[i] for i in ids) == text

For a known corpus, this tokenizer is simple. With a new character, however, it fails. You could add an unknown token, but then you lose the exact original text. Unicode matters too: one visible letter can consist of several code points. Different normalizations have to be chosen deliberately, not happen by accident.

Bytes as a robust foundation

UTF-8 turns any valid Unicode text into bytes between 0 and 255. A byte tokenizer needs only 256 regular tokens and a few special tokens. For it, there are no unknown regular bytes.

python
raw = 'Grüße 🐈'.encode('utf-8')
ids = list(raw)
text = bytes(ids).decode('utf-8')
assert text == 'Grüße 🐈'

An umlaut can need several bytes, and so can an emoji. Individual generated bytes are not necessarily valid UTF-8 text on their own. When an incomplete or corrupted generated sequence is decoded, your output has to handle that. Our tokenizer reconstructs valid input texts exactly; for arbitrary model outputs it uses replacement characters for invalid UTF-8 sequences.

Byte Pair Encoding step by step

In our project, BPE starts with bytes. It counts adjacent token pairs in the training corpus, picks the most frequent pair (on a tie, deterministically the one with the smallest IDs), gives its merged byte sequence a new ID and replaces the occurrences of that pair. Then it counts again. The procedure stops after a chosen number of merges or when no sufficiently frequent pair is left.

The word “hello” might first merge l,l into ll and later h,e into he. The exact merges come from the corpus; there is no prescribed linguistic split. In the browser experiment we use characters instead of bytes for readability, but the same pair-counting and replacement logic. This is explicitly a teaching run of character-level BPE.

A merge is not a replacement of every matching substring in the text. It acts on the current token sequence. For [a,a,a] and the merge (a,a), we replace from the left without overlap and get [aa,a]. If you replaced overlapping pairs at the same time, the result would not be well defined.

Vocabulary size is a trade-off

A larger vocabulary often shortens sequences. But it enlarges the embedding and output matrices and can create rare tokens that have few training examples. A smaller vocabulary needs fewer table parameters, but more positions for the same text. For our micro model, a few hundred tokens are practical; for a more serious small model, a larger tokenizer often makes sense.

Don't simply compare the token loss of two different tokenizers. A token can be a byte in one model and half a word in another. For such a comparison you need a common unit, such as bits per byte on identical texts, plus a clean way of measuring.

Special tokens are separate IDs

Our project reserves PAD, BOS, EOS, USER, ASSISTANT and SYSTEM. EOS marks the end of a document or an answer. The role markers belong to the chat formatting that comes later. Normal text is processed as bytes and cannot produce a reserved ID through some random string fragment. We must not automatically interpret a visible string like <assistant> in normal text as a role.

When you prepare the training data, you insert reserved IDs explicitly. On output you can skip them or show them as readable markers. “EOS” is not a word that turns up in the document by itself; it is a structural decision.

Train the tokenizer on training data only

First split the documents into training and validation. Then learn the merges on the training documents only. Otherwise your tokenizer receives information about held-out texts. This is a milder form of data leakage than training the model directly on the validation set, but it still changes the conditions of the experiment.

Save the bytes and the merges. The order of the merges is part of the encoder. For new text, our reference encoder applies the learned rules in that order. This naive procedure is slow but easy to check. It is a teaching tokenizer, not a replacement for highly optimized libraries in a production pipeline.

Check the round trip first

Test ASCII, umlauts, line breaks, emojis, empty strings and texts containing strings that look like special tokens. decode(encode(text)) == text must hold for valid input texts as long as you don't intend any normalization. Also check that saving and loading again gives the same encoder. Before a model learns anything, the way into the world of tokens and back should work.