How Large Language Models Are Trained
Pre-training on vast text, instruction tuning, and learning from human preferences — plus the scaling laws that made models grow.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. A chatbot like ChatGPT or Claude is not trained in one step. It goes through several stages, each teaching something different.
Three stages. First, pre-training: predict the next token across an enormous amount of text. Knowledge and language skills emerge. Second, instruction tuning on example conversations, so the model follows requests. Third, preference tuning, so answers become more helpful and safer.
Scaling laws. Researchers found scaling laws: as models, data and compute grow, the loss falls in a smooth and predictable way, with diminishing returns. That predictability justified training ever larger models. This curve is illustrative, but the pattern is real.
Learning from preferences. In preference tuning, people compare pairs of answers. Here, response A is clear and child friendly, so it is preferred. A reward model learns to predict these preferences, and the language model is tuned to earn higher reward, with methods like PPO, or more directly with DPO.
What each stage adds. Pre-training gives knowledge and language. Instruction tuning gives the ability to follow requests. Preference tuning shapes tone, helpfulness and safety. But none of them guarantees the truth, so important outputs still need checking.
Recap. To recap. Pre-train on vast text, instruction-tune on examples, preference-tune with human feedback, and scale up following predictable laws.