Tokenwerk · The LLM textbook

Chapter 33 · VI · Your own project · 31 minutes

Glossary, formulas and sources

A reference for the most important terms and checked primary sources for further reading.

Five formulas you should be able to explain

These short forms remind you of calculations you have already derived. The chapter links take you to a worked numerical example and a full explanation of every symbol.

Softmax: pi=exp⁡(zi)/∑jexp⁡(zj)p_i=\exp(z_i)/\sum_j\exp(z_j). It normalizes the scores over the possible next tokens. Step-by-step explanation.

Next-token loss: L=−log⁡ptargetL=-\log p_{target}. With several active targets, sum them and divide by their count. Step-by-step explanation.

Gradient step: w←w−η∇wLw\leftarrow w-\eta\nabla_w L. Eta is the learning rate; ∇ (nabla) denotes the list of derivatives, that is, the gradient. The arrow ← means: replace the old value with the result computed on the right. Step-by-step explanation.

Attention: softmax⁡(QK⊤/d+M)V\operatorname{softmax}(QK^\top/\sqrt d+M)V. Query and key provide the weights; values provide the information that gets mixed. M is the additive mask: 0 for allowed sources, minus infinity for forbidden ones. The attention calculation and masking.

KV cache: 2LBTHkvde2LBTH_{kv}de. The factors describe keys plus values for every layer, every batch example and every stored position. Calculation with sizes and units.

A good path through research papers

Start with the abstract, the problem statement and the architecture figure. Then write down in your own words: Which existing difficulty is the method meant to solve? Which part of the calculation changes? Which comparison conditions were used? Which limitations remain?

Only after that, try to recompute the central equation and write the smallest relevant piece of code. A paper with impressive performance numbers is not automatically a recipe for your small model. Transfer assumptions and experiments deliberately.

Foundations and small models

Modern architecture building blocks

Scaling and post-training

Official implementation documentation

  • PyTorch installation: the right installation for your operating system and backend.
  • SDPA: masks, dropout and kernel conditions of the attention function we use.
  • AdamW: optimizer arguments and implementation details.
  • Automatic Mixed Precision: autocast and gradient scaling.
  • MPS backend: Apple Silicon support and device selection.

The content of this textbook consists of learning examples explained independently, not translations of individual papers. Technical sources were looked up while it was written, on 6 October 2026. Library documentation may change later. The download describes the supported assumptions of its reference code; the links complement it.

Common misconceptions

"Attention automatically understands meaning." Attention is a learnable mechanism for mixing information. Which useful relationships emerge depends on training and on the overall architecture.

"More parameters mean more truth." The parameter count describes capacity, not fact-checking. Errors in the data and wrong output targets can persist in larger networks too.

"SFT is just formatting." SFT can also teach task knowledge and behavior. But it does not automatically make up for missing broad language competence, and with little data it can memorize heavily.

"A long context window is a long memory." A window defines the available input. Reliable use across large distances has to be trained and tested; persistent external storage is yet another function.

"A good loss proves a good product." The training metric is only part of the goal. An application also needs suitable task tests, error handling and system behavior.

All terms with examples

You can open the same explanations while reading, directly from the terms with a dotted underline. The full alphabetical collection is also included in the downloaded book.

Ablation

Evaluation tests performance; a benchmark is a fixed collection of test tasks. An ablation compares variants in which a single building block was changed on purpose.

Example: Compare two otherwise identical models with and without a particular normalization.

More in the chapter “Evaluating answers honestly”.

Derivative

The local rate of change of one quantity with respect to another. It describes how sensitively the result reacts to a very small change.

Example: Derivative −16: a small weight increase of 0.01 lowers the error by 0.16 in the local approximation.

More in the chapter “How numbers actually learn”.

AdamW

Optimization methods that keep statistics of past gradients for each parameter. AdamW separates an additional shrinking of the weights from the gradient-based step.

Example: A parameter with strongly fluctuating gradients gets a different scaling than one with steady gradients.

More in the chapter “Your first training loop”.

Activation

An intermediate value computed during the forward pass, often the output of an activation function. It arises for the current input and is not a permanent weight.

Example: A neuron has the sum −2. After ReLU, its activation is 0.

More in the chapter “From neurons to a small network”.

Activation function

A function that further processes the computed value of a neuron. Nonlinear activations allow more complex relationships.

Example: ReLU turns −3 into 0 and leaves 4 unchanged.

More in the chapter “Why neurons need an activation function”.

Attention

A calculation in which each position combines information from allowed positions using variable mixing weights.

Example: Values 10, 20, 30 with weights 0.5, 0.25, 0.25 give 17.5.

More in the chapter “Self-attention without the mystery”.

Output head

The final conversion of internal lists of numbers into the required output values.

Example: A position's list of 128 numbers becomes 320 logits for 320 possible tokens.

More in the chapter “The complete transformer block”.

Autograd

The automatic calculation of derivatives from the computation steps that were executed. PyTorch provides a system for this.

Example: After a forward pass, loss.backward() triggers the backward pass.

More in the chapter “Learning through several layers”.

Backpropagation

A method that traces the influences on the error backward through the individual calculations. It computes gradients; only the optimizer changes the weights.

Example: First determine the influence on the output, then on the intermediate value and finally on the earlier weight.

More in the chapter “Learning through several layers”.

Baseline

A baseline is a simple point of comparison. Our bigram baseline estimates the next token only from the token immediately before it.

Example: After a, b and c were each seen once: without smoothing, both get probability 1/2.

More in the chapter “Your first real language model”.

Batch

A group of examples processed together in one computation step. The batch size is the number of these examples.

Example: A batch of 4 text snippets has B = 4.

More in the chapter “Vectors, matrices and tensor shapes”.

Bias

An additional trainable term that is added regardless of the current input.

Example: In 3 × x + 1, the bias is 1. Even for x = 0, the output stays 1.

More in the chapter “An artificial neuron you can compute by hand”.

BPE

Byte Pair Encoding: a method that step by step merges frequent neighboring units. Byte-level BPE starts with bytes instead of characters.

Example: If "a" and "b" often appear next to each other, they can become a new unit "ab".

More in the chapter “From text to token IDs”.

Broadcasting

A rule that lets an operation reuse a smaller, compatible shape of numbers across larger axes.

Example: A bias with D entries is added to each of the many position lists with D entries.

More in the chapter “Vectors, matrices and tensor shapes”.

Byte

A byte stores 8 bits. UTF-8 is a fixed rule that translates text characters into one or more bytes.

Example: The ASCII character a needs one byte in UTF-8; many other characters need several.

More in the chapter “From text to token IDs”.

Checkpoint

A saved training state. To resume, you need not only the weights but also the settings and the optimizer state.

Example: Save after step 100 and later continue training from that state.

More in the chapter “Your Python workshop”.

Cross-entropy

The error measure for our next-token prediction: the negative logarithm of the probability of the correct token.

Example: A correct token with probability 0.5 gives an error of about 0.693.

More in the chapter “The learning goal: cross-entropy”.

Data parallelism

Ways to distribute training across devices. Data parallelism distributes examples, tensor parallelism distributes parts of calculations and pipeline parallelism distributes layers. Sharding splits up stored states.

Example: Two devices process different batches and synchronize their gradients.

More in the chapter “Scaling up without calculating blindly”.

Data leakage

Test information gets into training or model selection when it should not. This makes the measured performance look too good.

Example: The same text appears both in training and in validation.

More in the chapter “Data is part of the model”.

Decoder

A transformer processes lists of numbers with attention and small networks. Our decoder is a variant that reads only the text so far and predicts the next token.

Example: Embedding → several causal blocks → norm → logits for the continuation.

More in the chapter “The complete transformer block”.

Dimension

Depending on context, either the length of a list of numbers or one axis of an array of numbers. We say which meaning applies in the text.

Example: Embedding dimension 128 means: 128 numbers per token list. A tensor with shape [2,3,128] has three axes.

More in the chapter “Vectors, matrices and tensor shapes”.

Embedding

A trainable list of numbers that is read from a table by an ID.

Example: Token ID 1 selects row 1, for example [0.4, 0.6].

More in the chapter “Embeddings and shared representations”.

Entropy

A measure of how spread out a probability distribution is. A very confident choice has low entropy.

Example: Three equally likely tokens have higher entropy than [0.98, 0.01, 0.01].

More in the chapter “Language as a probability problem”.

EOS

Special kinds of tokens that mark boundaries and roles. BOS marks a beginning, EOS an end; USER and ASSISTANT mark roles in a conversation.

Example: An answer example can end with an EOS token after the last answer.

More in the chapter “From a text model to an answer model”.

Epoch

One complete pass through a fixed training dataset. When text windows are drawn at random, such a pass does not happen automatically.

Example: Once each of the 100 examples has been processed exactly once, an epoch is complete.

More in the chapter “Your first training loop”.

Epsilon

A very small positive safety constant, for example to keep a calculation from dividing by zero.

Example: LayerNorm adds epsilon to the variance before taking the square root.

More in the chapter “The complete transformer block”.

Exponential function

The function exp(x) = e to the power of x, with e approximately 2.718. It produces positive numbers and increases with x.

Example: exp(0) = 1, exp(1) approximately 2.718. Softmax uses these positive values.

More in the chapter “Language as a probability problem”.

Feature

A single entry of a representation. Learned features do not have to represent a property that humans can name.

Example: In [0.2, 0.7], 0.2 and 0.7 are the two components.

More in the chapter “Embeddings and shared representations”.

FlashAttention

FlashAttention computes dense attention block by block with less memory traffic. SDPA is PyTorch's interface for scaled dot-product attention. A kernel is the concrete, hardware-level compute routine.

Example: Depending on the device, an SDPA call may use a different compute routine; the function name does not guarantee any particular speedup.

More in the chapter “GQA, FlashAttention and the KV cache”.

FLOPs

A number of floating-point operations, that is, an amount of compute work. One PFLOP equals 10 to the power of 15 operations. Only FLOP/s describes a speed.

Example: 6 × parameters × training tokens is a rough estimate of the work for dense transformer training.

More in the chapter “Scaling up without calculating blindly”.

Forward pass

Computing a prediction from input to output with the current parameters.

Example: First h = 2 × w, then prediction = h × v. During this, w and v do not change yet.

More in the chapter “From neurons to a small network”.

GELU

A smooth activation function that softly dampens negative inputs instead of cutting them off hard at zero like ReLU.

Example: It changes the intermediate values between the two linear layers of the classic MLP.

More in the chapter “The complete transformer block”.

Generalization

Successfully applying learned patterns to examples that were not used before.

Example: After 1→3, 2→6 and 3→9, the model returns 12 for the new input 4.

More in the chapter “What does it mean for a program to learn?”.

Weight

A trainable multiplier. It determines how an input value contributes to a subsequent calculation.

Example: A weight of −2 turns input 3 into the contribution −6.

More in the chapter “An artificial neuron you can compute by hand”.

GQA

Variants of how query, key and value heads are distributed. GQA lets several query heads use shared key/value heads.

Example: 8 query heads and 2 key/value heads: every four queries share one K/V pair.

More in the chapter “GQA, FlashAttention and the KV cache”.

Gradient

The list of derivatives of the error with respect to the parameters. It describes each parameter's local influence.

Example: With a single weight, the gradient is just one number, say −16.

More in the chapter “How numbers actually learn”.

Gradient accumulation

Several small groups of examples are computed forward and backward one after another. Their gradients accumulate before a single shared parameter update.

Example: Four microbatches of 2 examples each can form one update step with effectively 8 examples.

More in the chapter “Your first training loop”.

A separate attention calculation with its own representations. Several heads can form different mixtures in parallel.

Example: With a total width of 128 and 4 heads, each head can compute with width 32.

More in the chapter “Causality and multiple heads”.

Hidden state

A list of numbers computed inside the network. It is an intermediate value, not a target specified by humans.

Example: The representation of a text position after the first transformer block is a hidden state.

More in the chapter “From neurons to a small network”.

Inference

Inference means applying a trained model. In autoregressive generation, one new token is appended at a time and the next prediction is made again.

Example: "The cat" first becomes "The cat sits" and then a longer sequence.

More in the chapter “From scores to text”.

Chain rule

A rule for derivatives of functions applied one after another: the local rates of change along a path are multiplied.

Example: If h reacts to w with factor 2 and the output reacts to h with factor 3, the combined influence is 6.

More in the chapter “Learning through several layers”.

Key

The list of numbers of a possible information source that is used for comparison.

Example: Query and key determine a score. The content that is mixed in afterward is in the value.

More in the chapter “Self-attention without the mystery”.

Corpus

The collection of texts from which training or test data is created.

Example: Ten separate stories form a small corpus.

More in the chapter “Data is part of the model”.

KV cache

A KV cache stores keys and values computed so far. Prefill reads the prompt; decode then adds new tokens one at a time.

Example: For the next token, earlier keys and values are reused instead of being recomputed from scratch.

More in the chapter “GQA, FlashAttention and the KV cache”.

LayerNorm

LayerNorm normalizes the components of a position and then learns a scale and a shift. Pre-norm means: normalization before the following operation.

Example: In a pre-norm block, x is normalized before attention is called.

More in the chapter “The complete transformer block”.

Learning rate

A setting for how large a learning step is. In the simple gradient step, it multiplies the derivative.

Example: With learning rate 0.05 and gradient −16, w = 1 becomes the new value 1.8.

More in the chapter “How numbers actually learn”.

Logarithm

Here the natural logarithm: the inverse function of exp. It asks which exponent produces a number.

Example: Because exp(0) = 1, log(1) = 0. For probabilities smaller than 1, log is negative.

More in the chapter “The learning goal: cross-entropy”.

Logit

An initially unbounded numerical value for a possible output. Only softmax turns all the logits together into probabilities.

Example: The logits [2,1,0] give approximately [0.665, 0.245, 0.090].

More in the chapter “Language as a probability problem”.

Loss

A number that rates how poorly the prediction fits the given target. By this measure, smaller is better.

Example: Prediction 2, target 6: the squared error is (2−6)² = 16.

More in the chapter “How do we measure a wrong answer?”.

Markov chain

A step-by-step random process whose next state depends only on the current state.

Example: The bigram model picks the next token from the row of the most recently generated token.

More in the chapter “Your first real language model”.

Mask

A rule that determines which entries may take part in a calculation. A causal mask forbids future text positions.

Example: Position 2 may read positions 1 and 2, but not 3.

More in the chapter “Causality and multiple heads”.

Matrix

A rectangular table of numbers with rows and columns.

Example: A 3-by-2 matrix contains six numbers.

More in the chapter “Vectors, matrices and tensor shapes”.

Matrix multiplication

A way of combining tables of numbers through dot products of their rows and columns. The inner sizes must match.

Example: [1,2] times [[3,4],[5,6]] gives [13,16].

More in the chapter “Vectors, matrices and tensor shapes”.

MiB

Binary units of memory: 1 MiB = 1024² bytes; 1 GiB = 1024³ bytes.

Example: 1,048,576 FP32 numbers at 4 bytes each take up 4 MiB.

More in the chapter “Precision and training memory”.

MLP

Multilayer perceptron: a network of successive linear layers with an activation in between. In the transformer, it works separately on each position.

Example: Width 8 → linear layer to 32 → activation → linear layer back to 8.

More in the chapter “The complete transformer block”.

Model

A fixed calculation rule together with its tuned parameters, which produces predictions from inputs.

Example: For prediction = input × w, the rule is a multiplication and w is the parameter.

More in the chapter “What does it mean for a program to learn?”.

MoE

An architecture with several subnetworks, of which only selected parts are used for each token.

Example: A routing mechanism sends a position's representation to a few out of many expert networks.

More in the chapter “What today's large systems need on top”.

MSE

Mean squared error: the average of the squared deviations between predictions and targets.

Example: Individual errors 1, 4 and 9 give MSE = 14/3.

More in the chapter “How do we measure a wrong answer?”.

Nats

The unit of information measures computed with the natural logarithm. It is not an additional model quantity.

Example: −log(0.5) gives approximately 0.693 nats. With a base-2 logarithm, the same information measure would be 1 bit.

More in the chapter “The learning goal: cross-entropy”.

Neuron

A computing unit that multiplies inputs by weights and adds up the contributions together with a bias. An activation function often follows.

Example: 2 × 1 + 1 × (−1) + 0.5 = 1.5 before the activation.

More in the chapter “An artificial neuron you can compute by hand”.

Neural network

A computational model made of connected units. Each unit processes numbers; adjustable weights determine the result.

Example: Two inputs are processed by two hidden neurons and then combined into one output.

More in the chapter “From neurons to a small network”.

Nonlinearity

A property of a calculation rule that cannot be described by a single weighted sum with a bias.

Example: ReLU has a kink. No single straight line can describe its entire curve.

More in the chapter “Why neurons need an activation function”.

Normalization

A calculation that adjusts the magnitude of values in a defined way. "Norm" can also mean the length of a vector; the context decides.

Example: LayerNorm subtracts the mean of a position's list and divides by its protected standard deviation.

More in the chapter “The complete transformer block”.

One-hot

A list with exactly one 1 and zeros everywhere else. The position of the 1 marks a selected category.

Example: For the second of three options: [0,1,0].

More in the chapter “Embeddings and shared representations”.

Optimizer

A rule that translates computed gradients into concrete parameter changes.

Example: A simple optimizer subtracts learning rate times gradient from the current weight.

More in the chapter “Your first training loop”.

Padding

Extra placeholders so that sequences of different lengths fit into one shared rectangular array of numbers. They should not count as real target text.

Example: A short sequence is filled up to the length of the longest sequence in the batch.

More in the chapter “The learning goal: cross-entropy”.

Parameter

A number in the model that can be changed during training. Input data and fixed settings are not trainable parameters.

Example: The weight w starts at 1 and is moved toward 3 during learning.

More in the chapter “What does it mean for a program to learn?”.

Perplexity

The exponential of the average per-token error. For equally likely options, it equals their number.

Example: Four equally likely tokens give a perplexity of 4.

More in the chapter “The learning goal: cross-entropy”.

Pretraining

The first broad training, in which the model predicts next tokens from many text sequences.

Example: Before question-answer training, the base model learns from general text.

More in the chapter “Your first training loop”.

Projection

Here, a learned linear conversion of a list of numbers into a new list, possibly with a bias.

Example: A linear layer converts a representation with 128 components into 320 logits.

More in the chapter “Embeddings and shared representations”.

Prompt

The text the model may read for the current prediction. Prefix means the beginning of a sequence given so far.

Example: For "The cat sits", this entire text so far is the context for the next token.

More in the chapter “From text to token IDs”.

Precision

The number format determines memory use, value range and rounding accuracy. Float32/FP32 uses 32 bits, FP16 and BF16 use 16 bits each.

Example: One million FP32 numbers need about 4 million bytes for their values alone.

More in the chapter “Precision and training memory”.

Quantization

Values are represented with fewer levels, that is, fewer bits. This saves memory but can increase rounding errors.

Example: A weight table is converted into 8-bit values plus scaling information.

More in the chapter “Precision and training memory”.

Query

The list of numbers of a reading position, compared with the keys of possible sources.

Example: Query [1,0] has a dot product of 1 with key [1,1].

More in the chapter “Self-attention without the mystery”.

RAG

Retrieval looks up suitable external information. RAG adds such retrieved text to the language model's input context.

Example: Before answering, relevant sections of a manual are retrieved and passed along.

More in the chapter “What today's large systems need on top”.

Computational graph

A representation of calculations as connected steps: the output of one step becomes the input of the next.

Example: Multiplication → second multiplication → deviation → square.

More in the chapter “Learning through several layers”.

Regularization

Methods that limit an overly specific fit to the available data. Additive smoothing adds small contributions to counts, including for unseen cases.

Example: Counts [0,1,1] become [1,2,2] after adding 1.

More in the chapter “Your first real language model”.

ReLU

A simple activation function: negative values become 0, all others are kept.

Example: ReLU(−2) = 0 and ReLU(2) = 2.

More in the chapter “Why neurons need an activation function”.

Residual connection

A direct addition path in which the original input is added to the result of an operation.

Example: [2, 5] + [0.3, −0.2] gives [2.3, 4.8].

More in the chapter “The complete transformer block”.

RLHF

Methods that use ratings or preferred answers for further training. RLHF uses a reward signal; DPO uses answer comparisons directly in a training objective.

Example: For the same question, answer A is preferred over answer B; this influences further training.

More in the chapter “What today's large systems need on top”.

RMSNorm

RMS is the root of the mean square. RMSNorm divides components by this value and learns a scale, without subtracting the mean first.

Example: [3,4] has RMS √12.5 ≈ 3.536.

More in the chapter “RMSNorm, RoPE and SwiGLU”.

RoPE

RoPE encodes positions by rotating pairs of components in queries and keys. Sine and cosine compute the coordinates of this rotation.

Example: A quarter turn turns [1,0] into [0,1]. The length stays the same.

More in the chapter “RMSNorm, RoPE and SwiGLU”.

Sampling

Sampling picks a token from a distribution. Greedy always takes the largest value. Top-k restricts the choice to k candidates; top-p to a set with a given total probability.

Example: With 70% sofa and 30% roof, random sampling can pick either. Greedy takes sofa.

More in the chapter “From scores to text”.

Layer

A group of computing units at the same stage of the network.

Example: Two neurons read the same inputs. Their two outputs form a hidden layer.

More in the chapter “From neurons to a small network”.

Seed

A starting number for a pseudorandom number generator. It helps repeat random decisions in the same environment.

Example: With the same seed, the same split into training and test documents can result.

More in the chapter “Data is part of the model”.

SFT

Further training on given inputs and desired outputs. Here, questions serve as context and answers as active targets.

Example: For the question "Why does ice melt?", a factually correct example answer is trained.

More in the chapter “From a text model to an answer model”.

Scalar

A single number, as opposed to a list or a table.

Example: The learning rate 0.01 is a scalar.

More in the chapter “Vectors, matrices and tensor shapes”.

Dot product

Multiply two lists of equal length component by component and add up the products. The result is a single number.

Example: [1,2] and [3,4] give 1×3 + 2×4 = 11.

More in the chapter “Vectors, matrices and tensor shapes”.

Softmax

A function that converts a list of numbers into positive values that sum to 1. Higher input numbers get higher probabilities.

Example: Three equal logits each give probability 1/3.

More in the chapter “Language as a probability problem”.

SwiGLU

SwiGLU mixes two learned paths of numbers by component-wise multiplication. One of them is first transformed by SiLU. SiLU(x) = x × sigmoid(x), sigmoid(x) = 1/(1+exp(−x)).

Example: For a gate value of 1 and a second value of 2, the result before the final projection is approximately 0.731 × 2 = 1.462.

More in the chapter “RMSNorm, RoPE and SwiGLU”.

Target

Target means the desired value. In next-token training, the target is shifted one position ahead of the input: the target shift.

Example: Inputs a, b, c have targets b, c, d.

More in the chapter “The learning goal: cross-entropy”.

Temperature

A positive number by which we divide the logits before softmax. It changes how strongly the choice concentrates on high scores.

Example: Below 1, the distribution usually becomes sharper; above 1, flatter. Equal logits stay equally likely.

More in the chapter “From scores to text”.

Tensor

An ordered array of numbers with one or more axes; even a single number can be stored as a tensor.

Example: [2,3,4] as a shape means 2 groups, each with 3 lists of 4 numbers.

More in the chapter “Vectors, matrices and tensor shapes”.

Token

A unit into which the tokenizer splits text. It can be a character, a piece of a word or a whole word.

Example: A tokenizer might split "learning" into "learn" and "ing". The actual split depends on the tokenizer.

More in the chapter “From text to token IDs”.

Tokenizer

The procedure for splitting text and translating the pieces into fixed numbers.

Example: "cat" is split into units; each unit gets its token ID.

More in the chapter “From text to token IDs”.

Training

Repeatedly computing predictions and changing parameters based on examples and errors.

Example: Forward pass → measure the error → compute gradients → change parameters.

More in the chapter “What does it mean for a program to learn?”.

Transpose

Swap the rows and columns of a matrix.

Example: A table with 2 rows and 3 columns gets 3 rows and 2 columns.

More in the chapter “Vectors, matrices and tensor shapes”.

Validation

Testing on separate data that is not used for the direct parameter updates. We use it to choose settings and training states.

Example: Training error falls, test error rises: the model may be adapting too closely to the training examples.

More in the chapter “Reading the loss and finding bugs”.

Value

The content list of a position, which attention mixes into the result according to the computed weights.

Example: A value of 20 with weight 0.25 contributes exactly 5 to the sum.

More in the chapter “Self-attention without the mystery”.

Variance

Variance is the average of the squared distances from the mean. The standard deviation is its square root.

Example: [1,3] has mean 2, variance 1 and standard deviation 1.

More in the chapter “The complete transformer block”.

Vector

An ordered list of numbers. The order of its entries matters.

Example: [0.2, 0.7] can be a two-component representation of a token.

More in the chapter “Vectors, matrices and tensor shapes”.

Distribution

The assignment of a probability to every possible alternative. All the probabilities sum to 1.

Example: sofa 0.5, roof 0.3, sea 0.2.

More in the chapter “Language as a probability problem”.

Vocabulary

The collection of all token types with their IDs. V is the number of these entries.

Example: With V = 320, each token ID can select one of 320 fixed units.

More in the chapter “From text to token IDs”.

Warmup

The initial phase in which the learning rate rises step by step to its intended value.

Example: Over the first 100 steps, it rises from a small value to the maximum learning rate.

More in the chapter “Your first training loop”.

Weight decay

A small additional shrinking of weights during training. It is meant to keep weights from growing too large.

Example: Alongside the actual learning step, a small fraction of the current weight is subtracted.

More in the chapter “Your first training loop”.

Weight tying

Several places in the computation use the same parameters. With weight tying, the input embedding and the output head share one weight table.

Example: One table with V × D numbers is used for the input lookup and for the output projection. A second such table is no longer needed.

More in the chapter “Embeddings and shared representations”.

Overfitting

The model fits its training examples better and better without getting correspondingly better on new examples.

Example: It completes learned sentences well but makes more errors on new sentences.

More in the chapter “Reading the loss and finding bugs”.