AI in Motion

Diffusion Models in Depth

Generative AIIntermediate1:386 chapters

The forward process adds noise on a schedule; a neural network learns to reverse it. See the real DDPM noise schedule and latent diffusion.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 What does the forward diffusion process do?
Q2 What is the network trained to predict?
Q3 Why does latent diffusion use a VAE?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. Diffusion models power most of today’s image generators. Their training has two halves: slowly destroying images with noise, and learning to undo it.

The forward process. The forward process adds a little Gaussian noise at each of a thousand steps, following a schedule. The green curve shows how much of the original image survives. At step 250 it is about 72 percent, at step 500 about 28 percent, and by the final step almost nothing: pure noise. Handily, we can jump straight to any step with one formula.

The reverse process. A neural network, usually a U-Net or a Transformer, is trained to predict the noise in a noisy image at any step. To generate, start from pure noise and repeatedly subtract the predicted noise, step by step, until an image appears.

Latent diffusion. Running diffusion on full-size pixels is expensive. Latent diffusion, used by Stable Diffusion, works on a small compressed version from a VAE. A text encoder turns the prompt into embeddings that guide every step, and the VAE decoder turns the result back into pixels.

Speed-ups. Early diffusion models needed hundreds of steps. Smarter samplers and distillation now produce good images in a handful of steps, and the same ideas generate video, audio and 3D shapes.

Recap. To recap. Add noise on a schedule, train a network to predict it, then reverse the process from pure noise. Latent diffusion does it in a compressed space, guided by text.