Tokenwerk · The LLM textbook

Chapter 29 · V · Modern language models · 4 minutes

Evaluating answers honestly

A small test battery beats the prettiest screenshot. This is how you measure what your model can do and where it fails.

The metric follows the goal

If your goal is simple text continuation, validation loss, grammatical clarity and coherence belong in the picture. If your goal is answering questions, factual correctness and fit to the task must be rated. A model can write good English and still answer the question wrongly.

You can report several criteria without squeezing them into a supposedly objective intelligence score. To get started, a small test battery fixed in advance and rated by hand is enough. That is limited, but verifiable.

Build the test questions before the final training run

Write at least 20 questions that you don't use for training or for choosing hyperparameters. Group them into simple facts, explanations, text editing, new phrasings and missing information. Note an expected answer or a rating sketch.

Example: the training data contains "Why does ice melt?". A test question could be "What happens to an ice cube if it sits in the warm kitchen for a long time?" This rephrasing tests more than word-for-word repetition, but its content stays closely related. For real generalization you also need new content or new combinations of tasks.

A simple rubric

For each answer, rate clarity, factual correctness and fit to the task with 0, 1 or 2. Note the specific error. The sum is a practical summary of this rubric, not a universal unit of measurement.

Criterion 0 1 2
Clarity Barely understandable Partly understandable Clear, coherent text
Correctness Wrong Partly right / significant gap Right for the task
Fit to the task Misses the point Only partly fitting Answers the specific question

The browser view lets you apply this rubric to a sample answer. It does not verify facts automatically. Such verification needs suitable references or a clearly defined, solvable checking problem.

Avoid training contamination

A test says little if its questions or nearly identical answers are in the training data. Check for exact and near matches. With large, freely available benchmarks, contamination is a serious issue; for your own small corpus, tracing things back by hand is often still possible.

Repeatedly tuning against the same test questions also turns them into a development set. For the final evaluation, hold back another set that you don't keep looking at.

Several samples instead of one lucky hit

If you use sampling, seeds produce differences. Either rate one fixed, uniform decoding method, or several seeds per prompt, and report the average as well as typical errors. You may deliberately test a best-of-k method, but then you have to state k and the selection rule.

A display that shows the eleventh attempt after ten failed ones does not reflect normal single-answer quality. It mixes model ability with selection effort.

Don't treat automatic self-evaluation as truth

A large model can help you create a rubric or find possible errors. But its rating is error-prone too and can mistake style for factual quality. For the core small set of test questions, you check the reference answers yourself.

For math or programming tasks, executable checkers are often better than a mere judgment. For a function, you can run tests. For a clear arithmetic problem, you can compare the number. Open-ended explanations leave more room for judgment.

Write a model card

Document the architecture, parameter count, tokenizer, training sources, number of tokens processed, known limitations and evaluation methods. Add examples of typical errors. This is not an advertisement for a model but a description of what it can be used for.

After the small demo run, our final model mainly demonstrates the pipeline. Only once larger data and separate evaluation show a benefit should you claim abilities beyond that. The effort spent on training is not a quality criterion in itself.

An export format for your tests

json
{"id":"ice-02","prompt":"What happens to ice in a warm kitchen?","reference":"It melts because heat is added.","category":"paraphrase"}

Save the answers with checkpoint name, seed, temperature and time. Then you can compare later variants on the same questions. Switching models without keeping the old outputs makes regressions hard to spot.

What you learn from errors

Does the model repeat patterns? Check sampling and data diversity. Does it ignore the question? Check the chat template and SFT examples. Is it clearly worded but wrong? Then it may lack the knowledge, or it transfers it poorly. These categories lead to different next experiments. "Train more" is only one possible response.