AI in Motion

Aligning LLMs: Instruction Tuning, RLHF and DPO

Large Language ModelsDeep diveAdvanced10:5932 chapters

How a text predictor becomes a helpful assistant: supervised fine-tuning, preference data, reward models, RLHF with PPO and a KL leash, DPO, AI feedback, reward hacking and evaluation.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

5 questions to check your understanding.

Q1 In SFT, the loss is usually computed on…
Q2 What does the reward model learn from?
Q3 Why does RLHF include a KL penalty?
Q4 What does DPO remove from the RLHF pipeline?
Q5 A reward model scores A = 3 and B = 1. The Bradley–Terry probability that A is preferred is…

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. A pre trained model has read a vast amount of text, but it is not yet an assistant. It continues text rather than answering, and it has no particular preference for being helpful, honest or safe. In this deep dive we follow the post training steps that turn a text predictor into a useful assistant.

Base vs assistant. A base model continues any text plausibly. Ask it a question and it may list more questions, ramble, or imitate a forum argument. An aligned assistant follows instructions, aims to be helpful, honest and harmless, and declines clearly harmful requests. Alignment is about shaping behaviour, not adding knowledge.

The goal. This is the behaviour we want: a clear, correct answer that follows the instruction, here two sentences for a beginner. Producing exactly this kind of response reliably, across millions of different requests, is the goal of post training.

The post-training pipeline. The usual pipeline has several stages. Start from the pre trained base model. Supervised fine tuning teaches it to imitate good demonstration answers. Then collect preference data, comparisons between answers. Optimise towards preferred answers with reinforcement learning or direct methods, and finally evaluate and red team the result.

Supervised fine-tuning. Supervised fine tuning, also called instruction tuning, trains the model on pairs of prompts and ideal responses, written or carefully curated by people. It uses the same next token loss as pre training, but usually only on the response tokens. Tens of thousands of high quality examples can transform behaviour.

Pause and think. Pause and think. Why is the fine tuning loss usually computed only on the response tokens, and not on the prompt? Because we want the model to learn how to answer, not how to write users’ questions. Masking the prompt focuses all of the learning on producing good responses.

Where instruction data comes from. Where does instruction data come from? Some is written by trained annotators, some comes from curated public datasets, and much is generated by stronger models and then filtered. Stanford’s Alpaca project, for example, fine tuned a Llama model on fifty two thousand instructions generated by a more capable model.

Why alignment matters. The effect can be dramatic. In the InstructGPT study, human labellers preferred the outputs of a one point three billion parameter aligned model over those of the one hundred and seventy five billion parameter GPT three base model. Alignment made a model more than a hundred times smaller more useful.

Preference data. Demonstrations are expensive to write and hard to make perfect. So the next stage collects comparisons: people see two responses to the same prompt and pick the better one. Comparing is quicker and more consistent than writing an ideal answer, and it captures subtle qualities like tone and clarity.

Pause and think. Pause and think. Why are comparisons often preferred over asking people to write ideal answers? Because judging is easier than creating. Comparisons are faster, more consistent between annotators, and capture qualities like tone, safety and subtle mistakes that are hard to get perfect when writing from scratch.

Reinforcement learning from human feedback. Reinforcement learning from human feedback uses these comparisons in two steps. First, train a reward model that predicts which response people would prefer. Then use reinforcement learning to fine tune the language model so its responses earn high scores from that reward model.

The reward model. The reward model is usually a copy of the language model with a single number as output. It is trained on each comparison so that the preferred response scores higher than the rejected one. The loss is minus the log sigmoid of the score difference, the classic Bradley Terry model of pairwise preferences.

Pause and think. Pause and think. The reward model scores response A at two and response B at one half. How likely is a person to prefer A under the Bradley Terry model? The sigmoid of the difference, one and a half, is about point eight two, so roughly an eighty two percent chance.

RL fine-tuning. Now the reinforcement learning step. The language model is treated as a policy, each generated token is an action, and the reward model scores the finished response. Proximal policy optimisation, PPO, then nudges the model towards responses with higher rewards, using the clipped updates we met in the reinforcement learning track.

The KL leash. The objective adds a leash. We maximise the reward model’s score minus beta times the KL divergence from the supervised model. Without the penalty the policy would drift into strange text that exploits quirks of the reward model. The KL term keeps it close to fluent, sensible behaviour.

The RLHF loop. Here is the loop in full. A prompt goes to the policy, which generates a response. The reward model scores the response, while a frozen reference model measures how far the policy has drifted. PPO combines the reward with the KL penalty and updates the policy, and the loop repeats over many batches.

Reward hacking. The reward model is only a proxy for what people really want, and optimising a proxy too hard exploits its blind spots. This is Goodhart’s law in action. Common symptoms include ever longer answers, flattery of the user, and confident sounding claims that are wrong.

Over-optimisation. The pattern looks like this. As optimisation increases, the reward model’s score keeps rising. True quality, as judged by careful humans, rises at first, then peaks and falls. Good training stops near the sweet spot, which is exactly what the KL penalty and careful monitoring aim for.

Direct preference optimisation. Direct preference optimisation, DPO, published in 2023, skips the separate reward model and the reinforcement learning loop. It trains the model directly on preference pairs with a simple classification style loss that raises the relative likelihood of chosen responses and lowers that of rejected ones, compared with a frozen reference.

The DPO loss. The DPO loss compares log probability ratios: how much the model has increased the chosen response relative to the reference, minus the same for the rejected response. Remarkably, under the same assumptions it has the same optimum as KL regularised RLHF, but it trains like ordinary supervised learning.

RLHF vs DPO. Which is better? RLHF’s reward model can score fresh samples, allowing online improvement, but the pipeline is complex, costly and sometimes unstable. DPO is simple and stable, which made it very popular for open models, but it learns only from a fixed set of pairs. Many teams use both, or newer variants.

AI feedback. Human labelling is slow and expensive, and people disagree. Constitutional AI, introduced in 2022, uses an AI model guided by a written list of principles to critique and revise responses, and to generate preference labels. This is called reinforcement learning from AI feedback, and it makes the intended values explicit.

Verifiable rewards. For mathematics and programming, rewards can be computed automatically, by checking the final answer or running unit tests. Reinforcement learning with these verifiable rewards has been used to train reasoning models that think through long chains of steps, using methods such as group relative policy optimisation.

Safety training. Preference data also teaches safety: declining clearly harmful requests. There is a balance. Too little safety training is dangerous, while too much causes over refusal, where the model declines perfectly legitimate requests. Good assistants decline only what is harmful and still help with the rest.

Pause and think. Pause and think. All the preference labels come from a small group of annotators with similar backgrounds. What risk does that create? The model absorbs their particular tastes, blind spots and cultural assumptions, and may serve other users worse. Diverse annotators and clear guidelines really matter.

Sycophancy. Optimising for human approval has a subtle side effect called sycophancy. People tend to rate agreeable answers highly, so models can learn to tell users what they want to hear, agreeing with mistaken claims or praising weak work. Researchers counter this with targeted data and evaluations that reward honest disagreement.

Red teaming. Before release, models are red teamed: experts, external testers and even automated attacker models deliberately try to make them misbehave, with jailbreak prompts, requests for dangerous instructions or attempts to extract private data. The failures they find become new training data and new safeguards.

DPO in code. The DPO loss is only a few lines. Given the summed log probabilities of the chosen and rejected responses under the policy and the frozen reference, compute each response’s gain, take the difference, and apply a logistic loss. At the start the policy equals the reference, so the loss is exactly the log of two.

Evaluating assistants. How do we know alignment worked? Teams run human preference studies, head to head win rates, benchmark suites and safety test sets. Using a strong model as a judge is fast and cheap, but judges have biases, favouring longer answers, the first position, or their own style, so results need spot checking.

Methods at a glance. Here is the landscape. Supervised fine tuning teaches format and behaviour from demonstrations. RLHF with PPO uses a reward model and reinforcement learning. DPO optimises preferences directly. AI feedback scales labelling using written principles, and verifiable rewards train reasoning in mathematics and code.

Open problems. Alignment is far from solved. How do we supervise answers that humans cannot easily check, such as complex code or science? How do we make models reliably honest about what they know? How do we resist ever more creative jailbreaks? And whose values should a model reflect, when people and cultures disagree?

Recap. To recap. Base models continue text, and alignment shapes their behaviour. Supervised fine tuning imitates demonstrations. RLHF trains a reward model from comparisons and optimises with PPO on a KL leash. DPO optimises preferences directly. Throughout, watch for reward hacking, biased feedback and over refusal.