AI in Motion

How Large Language Models Are Trained

Generative AIIntermediate1:266 chapters

Pre-training on vast text, instruction tuning, and learning from human preferences — plus the scaling laws that made models grow.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 What is the objective during pre-training?
Q2 What does a reward model learn in RLHF?
Q3 What do scaling laws describe?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. A chatbot like ChatGPT or Claude is not trained in one step. It goes through several stages, each teaching something different.

Three stages. First, pre-training: predict the next token across an enormous amount of text. Knowledge and language skills emerge. Second, instruction tuning on example conversations, so the model follows requests. Third, preference tuning, so answers become more helpful and safer.

Scaling laws. Researchers found scaling laws: as models, data and compute grow, the loss falls in a smooth and predictable way, with diminishing returns. That predictability justified training ever larger models. This curve is illustrative, but the pattern is real.

Learning from preferences. In preference tuning, people compare pairs of answers. Here, response A is clear and child friendly, so it is preferred. A reward model learns to predict these preferences, and the language model is tuned to earn higher reward, with methods like PPO, or more directly with DPO.

What each stage adds. Pre-training gives knowledge and language. Instruction tuning gives the ability to follow requests. Preference tuning shapes tone, helpfulness and safety. But none of them guarantees the truth, so important outputs still need checking.

Recap. To recap. Pre-train on vast text, instruction-tune on examples, preference-tune with human feedback, and scale up following predictable laws.