Classifier-Free Guidance
How image generators follow prompts more closely: combine a conditional and an unconditional prediction and push further in the prompt’s direction.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. Image generators have a setting called guidance scale or CFG. It controls how strongly the image follows your prompt. Here is what it does under the hood.
Vector arithmetic. At each step, the model predicts the noise twice: once without the prompt, the grey arrow, and once with the prompt, the blue arrow. The difference between them points in the prompt’s direction. Guidance takes the unconditional prediction and moves w times along that direction. With w equal to one we get the plain conditional prediction. With w of three or seven and a half, we push further, so the prompt has a stronger effect.
Training trick. The trick is in training: the prompt is randomly dropped some of the time, so a single network learns to predict both with and without the prompt. No separate classifier is needed, which is why it is called classifier free.
Choosing w. Low guidance gives varied, natural images that may drift from the prompt. High guidance follows the prompt closely, but too high produces harsh, over-saturated images. Values around seven are a common default.
Recap. To recap. Guidance combines two predictions and pushes in the prompt’s direction. Prompt dropout during training makes it possible, and the scale trades prompt adherence for variety.