AI in Motion

Generative AIAdvanced1:19 video5 chapters

Classifier-Free Guidance — lecture notes

How image generators follow prompts more closely: combine a conditional and an unconditional prediction and push further in the prompt’s direction.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Classifier-Free Guidance

Image generators have a setting called guidance scale or CFG. It controls how strongly the image follows your prompt. Here is what it does under the hood.

0:122. Vector arithmetic

Vector arithmetic — Classifier-Free Guidance

At each step, the model predicts the noise twice: once without the prompt, the grey arrow, and once with the prompt, the blue arrow. The difference between them points in the prompt’s direction. Guidance takes the unconditional prediction and moves w times along that direction. With w equal to one we get the plain conditional prediction. With w of three or seven and a half, we push further, so the prompt has a stronger effect.

0:363. Training trick

Training trick — Classifier-Free Guidance

The trick is in training: the prompt is randomly dropped some of the time, so a single network learns to predict both with and without the prompt. No separate classifier is needed, which is why it is called classifier free.

0:534. Choosing w

Choosing w — Classifier-Free Guidance

Low guidance gives varied, natural images that may drift from the prompt. High guidance follows the prompt closely, but too high produces harsh, over-saturated images. Values around seven are a common default.

1:075. Recap

Recap — Classifier-Free Guidance

To recap. Guidance combines two predictions and pushes in the prompt’s direction. Prompt dropout during training makes it possible, and the scale trades prompt adherence for variety.

Key takeaways

  • Classifier-free guidance combines conditional and unconditional predictions.
  • ε̂ = ε_uncond + w · (ε_cond − ε_uncond); w = 1 is plain conditioning.
  • Randomly dropping the prompt during training teaches both predictions.
  • Higher guidance increases prompt adherence but reduces diversity.

Check yourself

  1. What happens when w = 1?
    Show answer

    You get the ordinary conditional prediction — ε_uncond + 1·(ε_cond − ε_uncond) = ε_cond.

  2. How does one network learn the unconditional prediction?
    Show answer

    The prompt is randomly dropped during training — Prompt dropout teaches both modes.

  3. What is a downside of very high guidance?
    Show answer

    Over-saturated, less varied images — Pushing too far distorts images.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/classifier-free-guidance.html