Tokenwerk · The LLM textbook

Chapter 31 · V · Modern language models · 5 minutes

What today's large systems need on top

MoE, long contexts, multimodality and post-training. A map for your next step instead of a list of magic terms.

The decoder is only one part of the system

A capable assistant system combines model weights with data selection, post-training, evaluation, an inference engine, tool access and product logic. Our course project implements the inner path of the language model. It does not implement a research engine, permanently reliable fact checking or a complete production service.

This boundary is nothing to be embarrassed about: it helps you attribute causes correctly. If a system can read a current web page, that information doesn't have to be stored in its weights. If it uses a calculator, a correct multiplication is not necessarily just an internally learned answer.

Mixture of experts

MoE typically replaces certain dense feed-forward layers with several expert networks plus a router. For each token, the router picks a small subset of experts. This allows many parameters to exist while only some of them are activated per token.

Total parameters describe, among other things, the memory for all expert weights. Active parameters are closer to the compute path actually used per token, but they are not a complete description of costs either: routing, attention and communication come on top. An MoE model with many total parameters still has to keep all the weights it needs available somewhere.

MoE is now standard in many large open models, for example Mixtral (8 experts, 2 active per token, around 47 billion total and 13 billion active parameters) or DeepSeek-V3 (671 billion total, 37 billion active per token, with many small and shared experts). For our course, the general router/expert idea is what matters, not a claimed current ranking of all models. Mixtral of Experts.

Why MoE shouldn't be your first step

A router can overuse certain experts or barely use others. Load balancing, capacity limits, extra objective terms and distributed communication make training more complex. For a small local learning model, a dense decoder can be easier to understand and even more practical.

If you build a mini MoE later, first check how many tokens each expert receives and whether all of them train. A falling overall loss alone does not show that the router is working well. A teaching extension can start with two to four small experts.

Long contexts and sparse attention

A larger configured length is not enough. The model has to learn to use relevant information across such distances. Dense attention becomes computationally expensive; sparse methods leave out certain connections, such as by using local windows or selected global positions. That can reduce costs, and it changes the path information takes.

So FlashAttention and sparse attention are not the same thing. And a large context window is no proof that the model reliably finds a relevant fact at every distance. Evaluate long contexts with controlled positions, distractions and tasks that actually need the information.

Multimodality

A multimodal model needs a path from images, audio or video to internal representations. That can happen through separate encoders and projection modules, or through other token-like representations. A text decoder doesn't become visual just because you call a token "image".

You need suitable data, a learning objective and an interface between modalities. For your first fully self-built text decoder, it pays to leave this extension for later on purpose.

Preference training

SFT trains on desired answers. Preference data often contains the same input with one preferred and one less preferred answer. The goal is to make the preferred output relatively more likely. A penalty term (parameter β) keeps the model close to a frozen reference model, usually the SFT checkpoint.

DPO is a method that uses a direct objective on such pairs instead of requiring a full classic RLHF pipeline with a separately trained reward model and PPO steps. But it doesn't remove the question of whether the preference data is factually good and fits the goal. Direct Preference Optimization.

RL with verifiable rewards

For some tasks, results can be checked automatically, such as passing program tests or correct numbers. An RL method can generate answers, score their reward and adjust the policy. What matters is that the reward measures the actual goal. A faulty checker can train the model to find workarounds instead of solutions.

Since 2024/25, leading systems have used such verifiable rewards to train so-called reasoning models: before answering, the model generates a long chain of thought. More compute at inference time – longer "thinking" – can substantially raise the success rate on math and programming tasks; this is called test-time compute. A widely used method is GRPO, which compares several answers to the same question with each other instead of using a separate value model. A well-known open example is DeepSeek-R1. Here too: the chain of thought is generated text, not a guaranteed view into the internal computation.

A reward is not simply another text corpus. Sampling, variance, credit assignment and regularization make the optimization different from ordinary SFT. Build a good evaluation first, before you experiment with complex RL.

Tools and retrieval

Retrieval brings relevant external text passages into the context. A tool call can calculate or query a database. The model has to choose these abilities appropriately and use the results correctly. This can make the overall system more capable without the decoder alone mastering every task.

For your own model, a pointless tool call made too early is not progress. Check the input format, the result format and failure cases. System evaluation also covers whether tools are used at the right moment.

What "modern" really requires

There is no universal architecture checklist. Good systems choose building blocks to suit their data, budget and use, and they demonstrate their effect. So your next step should be a precise question: "Can a mini MoE improve my validation loss with the same active compute budget?" can be investigated. "I'll add everything modern" is not an experiment.