How language models are pretrained
Understand the data, objective, scaling and stability work behind large language model pretraining
A taste of a lesson
Roughly how much compute does it take to train a 7 billion parameter model on 2 trillion tokens?
Use the rule of thumb of about 6 operations per parameter per token: 6 times 7 times 10 to the 9 times 2 times 10 to the 12 gives about 8.4 times 10 to the 22 floating point operations. To turn that into time, divide by what your hardware sustains in practice, which is usually well below its peak figure because of communication and memory overheads. Treat the result as an order of magnitude estimate, not a budget. Your turn: if you doubled the tokens but halved the parameters, how would the compute change?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain what next token prediction teaches a model and what it does not
- Describe a pretraining data pipeline including filtering, deduplication and mixing
- Estimate training compute from parameter count and token count
- Interpret scaling trends and the trade off between training and inference cost
- Distinguish what pretraining provides from what post training adds
Lesson plan
- 1 The next token objective Understand the pretraining loss and what kinds of knowledge it rewards. Start
- 2 Building the dataset Follow raw text through extraction, filtering and deduplication. Start
- 3 Tokenizers and data mixtures See how tokenizer choice and source weighting shape the model. Start
- 4 Compute and scaling trends Estimate training compute and interpret how loss scales with model and data size. Start
- 5 Keeping a long run stable Handle warmup, loss spikes, precision and checkpoints in large training runs. Start
- 6 From base model to assistant Separate what pretraining provides from what instruction and preference tuning add. Start
Try asking
About this tutor
For practitioners who know transformers and want to understand what actually happens when a large language model is pretrained. You will start from the next token prediction objective and what it does and does not teach, then follow the real work: collecting and filtering text, deduplication, choosing a tokenizer, mixing data sources, and checking for test set contamination. Then you study scaling: how loss improves with compute, parameters and data, what compute optimal means in general terms, and why many models are trained on more data than that. The final lessons cover training stability, mixed precision and what pretraining gives compared with later post training stages.
Reviews
4.7
3 ratingsSample
- Mohammed J.Sample
I liked that scaling laws were presented as empirical trends with limits, not magic. The stability lesson matched problems I had seen in smaller runs.
- Shreya B.Sample
Strong on data quality and contamination. I wish there was more on tokenizer training, but the tutor gave good pointers when I asked.
- Anders L.Sample
The 6ND estimate and the inference cost discussion explained why our team picked a smaller model trained longer. Very clear lessons on data pipelines too.
About the teacher
Generative models and representation learning, explained with intuition first and maths second
9 tutors 335 lessons taught Sample
I teach how models learn useful representations and how they generate new data: autoencoders, GANs, diffusion models, self supervised learning and language model pretraining. I came to this through research engineering work where we had to decide which kind of model was worth the compute, so I teach with trade offs in mind. Each topic starts with a picture or...
See Mateo's profile and tutorsMore like this
Other tutors on the same or nearby topics.