Three resources belong together
A larger model can represent more patterns, but it needs suitable data and enough optimization. More data only helps if the model gets enough compute to use it. A higher step count on the same small corpus increases the token budget processed, but not the amount of new information.
A floating-point operation is a single calculation on numbers stored in that format, such as an addition or a multiplication. FLOPs counts such operations. The term describes an amount of work; time only comes in when we divide by operations per second.
That's why we note the parameter count N, the training tokens D actually processed, and the amount of work. For a dense transformer, is a commonly used rough approximation of pretraining FLOPs. It simplifies several details, in particular context-dependent attention costs. The 6 roughly breaks down into 2 FLOPs per parameter and token in the forward pass and 4 in the backward pass (Kaplan et al., 2020; there excluding embedding parameters). N counts parameters, and in this chapter D counts training tokens processed; so here D explicitly does not mean the embedding width. The estimate is not a universal law of performance.
For 15 million parameters and 300 million training tokens, the approximation gives FLOPs, or 27 PFLOP of compute work. Here PFLOP is an amount of operations. PFLOP/s would be a speed.
From compute budget to time
Divide the work required by a throughput you actually achieve, not by a marketing peak figure. A theoretical GPU performance figure assumes ideal shapes, dtypes and utilization. Data processing, startup costs, communication and small operations can lower real throughput considerably.
For your project, directly measuring training tokens per second is often more practical. If you measure 2,000 tokens per second, 10 million processed tokens take roughly 5,000 seconds, not counting extra pauses, evaluation and saving time. This number is a calculation from an assumed measurement, not a claim about your Mac.
The idea behind Chinchilla
The Chinchilla paper studied how a fixed pretraining compute budget can be split between model size and training data. Its results motivate scaling both together. A commonly cited rough rule of thumb of about 20 training tokens per parameter is tied to specific assumptions and a compute-optimal training point. It is not a quality minimum and not a law for every model.
If you later want to run a small model very often, training a smaller model for longer can make economic sense, even though it doesn't reach the same loss with the minimum pretraining effort. Training costs and inference costs are different optimization goals. In practice, smaller models are therefore often trained far beyond 20 tokens per parameter: Llama 3 8B saw over 15 trillion tokens, almost 2,000 tokens per parameter. Training Compute-Optimal Large Language Models.
Memory is a different bottleneck
A model can fit in memory and still be too slow for your time frame. It can also be fast enough computationally but not fit with the batch and context you want. Parameter count alone determines neither runtime nor peak memory use.
First enlarge just one variable: width, depth or context. Then measure token throughput, full training memory and validation loss. That way you can tell whether your bottleneck is matrix throughput, memory, Python overhead or data preparation.
Distributed training: an overview
Data parallelism replicates the model and distributes different batches; gradients are synchronized. Tensor parallelism splits parts of individual operations across devices. Pipeline parallelism distributes consecutive layers. Sharding the optimizer and parameters can reduce memory per device.
These methods add communication, synchronization and failure cases. An architecture with millions of parameters on a single device is therefore a sensible first step. Distributed systems should solve a measured problem, not just look more modern.
Scale data preparation too
The naive byte-level BPE in our download is easy to understand but not efficient for large data. Long Python lists and repeated merge counts can become the bottleneck. A productive next step is to keep the format and the tests and replace the tokenizer trainer with an optimized implementation.
Likewise, memory mapping, parallel data readers and pre-tokenized shards can help. Don't silently change the tokenizer or the validation method while you do this. A faster pipeline that sees different data is not a pure performance comparison.
Budget before the run
First estimate data tokens, model parameters, memory and time from a short benchmark. Then do a small pilot run. If the loss doesn't respond sensibly, don't start the ten-times-longer run yet. The most expensive experiment is often a long training run with an error that was already visible in the first batch.
Three sensible milestones
First: a correct micro model can overfit a small batch. Next: a small model beats the bigram baseline on independent documents. Finally: a larger suitable corpus plus SFT delivers understandable and partly reliable answers on new questions fixed in advance. These milestones build ability and evidence together.