AI in Motion

Generative AIIntermediate1:26 video6 chapters

How Large Language Models Are Trained — lecture notes

Pre-training on vast text, instruction tuning, and learning from human preferences — plus the scaling laws that made models grow.

▶ Watch the animated lecture

0:001. Introduction

Introduction — How Large Language Models Are Trained

A chatbot like ChatGPT or Claude is not trained in one step. It goes through several stages, each teaching something different.

0:092. Three stages

Three stages — How Large Language Models Are Trained

First, pre-training: predict the next token across an enormous amount of text. Knowledge and language skills emerge. Second, instruction tuning on example conversations, so the model follows requests. Third, preference tuning, so answers become more helpful and safer.

0:263. Scaling laws

Scaling laws — How Large Language Models Are Trained

Researchers found scaling laws: as models, data and compute grow, the loss falls in a smooth and predictable way, with diminishing returns. That predictability justified training ever larger models. This curve is illustrative, but the pattern is real.

0:424. Learning from preferences

Learning from preferences — How Large Language Models Are Trained

In preference tuning, people compare pairs of answers. Here, response A is clear and child friendly, so it is preferred. A reward model learns to predict these preferences, and the language model is tuned to earn higher reward, with methods like PPO, or more directly with DPO.

1:025. What each stage adds

What each stage adds — How Large Language Models Are Trained

Pre-training gives knowledge and language. Instruction tuning gives the ability to follow requests. Preference tuning shapes tone, helpfulness and safety. But none of them guarantees the truth, so important outputs still need checking.

1:176. Recap

Recap — How Large Language Models Are Trained

To recap. Pre-train on vast text, instruction-tune on examples, preference-tune with human feedback, and scale up following predictable laws.

Key takeaways

  • LLMs are pre-trained on next-token prediction over very large text corpora.
  • Instruction tuning teaches them to follow requests.
  • RLHF or DPO uses preference comparisons to improve helpfulness and safety.
  • Scaling laws show loss falling predictably with model size, data and compute.

Check yourself

  1. What is the objective during pre-training?
    Show answer

    Predict the next token — Pre-training is next-token prediction on text.

  2. What does a reward model learn in RLHF?
    Show answer

    To predict which answers people prefer — It is trained on human preference comparisons.

  3. What do scaling laws describe?
    Show answer

    How loss falls predictably as size, data and compute grow — Performance improves smoothly with scale.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/how-llms-are-trained.html