A token ID is at first only a row number
Imagine a table with three rows of two numbers each:
| ID | Token | First component | Second component |
|---|---|---|---|
| 0 | cat | 0.2 | 0.7 |
| 1 | dog | 0.4 | 0.6 |
| 2 | sofa | -0.3 | 0.1 |
For ID 1 we look up [0.4, 0.6]. This list of numbers is called an embedding. The numbers here are just a made-up example to calculate with. In a real model they start as small random values and are changed during training. We do not decide by hand that the first number should mean “kind of animal” or the second “size”.
The ID itself is not a measurement: with ID 1, dog is neither bigger nor better than cat with ID 0. Only the trainable list of numbers creates a representation the network can sensibly keep computing with. Its length is called the embedding dimension or width D, here D = 2.
A one-hot vector would be a list with a 1 only at the ID in question and zeros everywhere else: for ID 1, that is [0,1,0]. Multiplying such a selection by the table gives the same row. In practice we look the row up directly and save ourselves this long list of zeros.
The embedding table
An embedding table contains a vector with components for every token. It has shape . For ID 42, the model reads row 42. This is not a search over meanings but an index lookup. The vectors start out random and are adjusted during training.
import torch
from torch import nn
embedding = nn.Embedding(100, 16)
ids = torch.tensor([[4, 12, 4]])
x = embedding(ids)
assert x.shape == (1, 3, 16)
assert torch.equal(x[0, 0], x[0, 2])At this point, the two occurrences of ID 4 give the same vector. Only later do positions and context give them different representations.
Why IDs have no semantic order
Imagine “dog” has ID 10 and “cat” ID 900. The numerical distance of 890 is meaningless. The distance between their learned vectors, on the other hand, can become relevant for certain tasks. An embedding replaces a discrete address with a trainable list of numbers. “Discrete” here means: one of a fixed set of IDs. The entries of the list, by contrast, can take in-between values such as 0.2 or 0.21. With these lists we can carry out the multiplications and additions we already know.
You could write this lookup as multiplying a one-hot vector by a table. A one-hot vector has a one in exactly one position. The product selects exactly one row of the table. The actual implementation skips the huge, mostly zero vector and uses the index directly.
No predefined drawers of meaning
Dimension 0 does not automatically mean “animal”, dimension 1 “colour” and dimension 2 “friendliness”. Different coordinate systems can implement the same function if the transformations that follow are adjusted. Sometimes directions can be interpreted; that does not mean every entry encodes a fixed human property.
Similar contexts can encourage similar representations, but “semantically similar” is not a guaranteed result for every layer and every distance metric. A small byte model, for example, first learns patterns of characters and word pieces. We don't give it a table of relationships between all concepts.
A preview: the same ID in a different context
The input embedding of “bank” stays the same for the same token ID. In a later network, this list is processed further together with the preceding text. In “money in the bank” and “sitting on the river bank”, this can lead to different internal representations. We call such an intermediate value a hidden state, an internal state.
How the network takes earlier text into account is explained only in the attention chapter. For this chapter it is enough to know: the initial table lookup is always the same; the processing that follows can take context into account.
A projection to the next distribution
Suppose further processing gives the list h = [1, 2]. For three possible tokens we use the rows [1, 0], [0, 1] and [1, 1] of a weight table. Their dot products with h give 1, 2 and 3. Without a bias, our logits are therefore z = [1, 2, 3]. Softmax could turn these into a distribution next.
In general we write . h is the input list with D components. W_out is a table with V rows and D columns. Each row produces one logit; b is a bias list with V entries that is added afterwards. z is the finished list of V logits. “out” stands for output.
In technical texts you will come across the notation for this. You read the symbol ∈ as “is an element of”, and ℝ denotes the real numbers. Together, here it just means: h is a list with D real-valued components. correspondingly denotes the list length V of the output. In the computer, these numbers are approximated by a finite number format.
With weight tying, the input embedding and the output projection share the same parameter matrix. In PyTorch it looks like this:
self.embed = nn.Embedding(V, D)
self.head = nn.Linear(D, V, bias=False)
self.head.weight = self.embed.weightThe shared matrix is trained in both roles. This saves parameters and makes the two roles structurally related. It is no guarantee that every architecture gets better with tying. In the course project, it keeps small models manageable.
Going deeper: measuring the similarity of two lists
Compare a = [1, 0] with b = [2, 0]. Both point in the same direction, but b is twice as long. The dot product is 2. We compute the length of a vector as the square root of the sum of its squared components: a has length √(1²+0²) = 1, b has length √(2²+0²) = 2.
If we divide the dot product by both lengths, we get 2/(1×2) = 1. This cosine similarity is 1 for the same direction, −1 for opposite directions and 0 for directions at right angles. You don't need to compute any angles for it at first.
a and b are the vectors being compared. The dot means dot product; the double vertical bars mean the length of each vector. The denominator multiplies both lengths. If a list has length zero, this division is not defined. The experiment then shows no valid similarity value.
Another measure is the Euclidean distance, the ordinary distance: take the differences per component, square them, add them up and take the square root. The distance from [1, 0] to [2, 0] is 1. So both lists have the same direction, but not the same position.
The experiment uses freely adjustable two-dimensional lists. They are not learned embeddings of a real language model. Even similar learned vectors do not automatically prove the same meaning; that depends on the training and the measurement method.
What embeddings have to learn
With the embedding table, all occurrences of an ID share parameters. The gradient of a training step can therefore carry information from many positions into the same row. Other parts of the model are shared as well: the same transformer block processes every time position, rather than a separate network per position.
This is a fundamental advantage over a collection of independent rules. Parameter sharing links examples together, but it also demands compromises: a word piece can occur with many meanings. The context path later has to tell these cases apart.
A small parameter calculation
With and , the table contains 1,048,576 parameters. In float32 that is about four MiB of weight data. Two tables would be eight MiB. The memory needed for training is larger, because derivatives and optimizer states are added. So don't reason “the embedding fits in RAM, so training fits”.