Classifier-Free Guidance — lecture notes
How image generators follow prompts more closely: combine a conditional and an unconditional prediction and push further in the prompt’s direction.
0:001. Introduction

Image generators have a setting called guidance scale or CFG. It controls how strongly the image follows your prompt. Here is what it does under the hood.
0:122. Vector arithmetic

At each step, the model predicts the noise twice: once without the prompt, the grey arrow, and once with the prompt, the blue arrow. The difference between them points in the prompt’s direction. Guidance takes the unconditional prediction and moves w times along that direction. With w equal to one we get the plain conditional prediction. With w of three or seven and a half, we push further, so the prompt has a stronger effect.
0:363. Training trick

The trick is in training: the prompt is randomly dropped some of the time, so a single network learns to predict both with and without the prompt. No separate classifier is needed, which is why it is called classifier free.
0:534. Choosing w

Low guidance gives varied, natural images that may drift from the prompt. High guidance follows the prompt closely, but too high produces harsh, over-saturated images. Values around seven are a common default.
1:075. Recap

To recap. Guidance combines two predictions and pushes in the prompt’s direction. Prompt dropout during training makes it possible, and the scale trades prompt adherence for variety.
Key takeaways
- Classifier-free guidance combines conditional and unconditional predictions.
- ε̂ = ε_uncond + w · (ε_cond − ε_uncond); w = 1 is plain conditioning.
- Randomly dropping the prompt during training teaches both predictions.
- Higher guidance increases prompt adherence but reduces diversity.
Check yourself
- What happens when w = 1?
Show answer
You get the ordinary conditional prediction — ε_uncond + 1·(ε_cond − ε_uncond) = ε_cond.
- How does one network learn the unconditional prediction?
Show answer
The prompt is randomly dropped during training — Prompt dropout teaches both modes.
- What is a downside of very high guidance?
Show answer
Over-saturated, less varied images — Pushing too far distorts images.
Go deeper
- Guidance in Diffusion Models: Classifier and Classifier-Free Guidance · The AI Lecture Hall
- Diffusion Models: Generating by Learning to Denoise · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/classifier-free-guidance.html