Diffusion Models in Depth — lecture notes
The forward process adds noise on a schedule; a neural network learns to reverse it. See the real DDPM noise schedule and latent diffusion.
0:001. Introduction

Diffusion models power most of today’s image generators. Their training has two halves: slowly destroying images with noise, and learning to undo it.
0:102. The forward process

The forward process adds a little Gaussian noise at each of a thousand steps, following a schedule. The green curve shows how much of the original image survives. At step 250 it is about 72 percent, at step 500 about 28 percent, and by the final step almost nothing: pure noise. Handily, we can jump straight to any step with one formula.
0:343. The reverse process

A neural network, usually a U-Net or a Transformer, is trained to predict the noise in a noisy image at any step. To generate, start from pure noise and repeatedly subtract the predicted noise, step by step, until an image appears.
0:524. Latent diffusion

Running diffusion on full-size pixels is expensive. Latent diffusion, used by Stable Diffusion, works on a small compressed version from a VAE. A text encoder turns the prompt into embeddings that guide every step, and the VAE decoder turns the result back into pixels.
1:115. Speed-ups

Early diffusion models needed hundreds of steps. Smarter samplers and distillation now produce good images in a handful of steps, and the same ideas generate video, audio and 3D shapes.
1:246. Recap

To recap. Add noise on a schedule, train a network to predict it, then reverse the process from pure noise. Latent diffusion does it in a compressed space, guided by text.
Key takeaways
- The forward process adds Gaussian noise following a schedule (e.g. linear β, 1000 steps).
- In the demo, √ᾱ is about 0.72 at t = 250 and 0.28 at t = 500.
- A network learns to predict the noise; generation reverses the process.
- Latent diffusion runs in a VAE latent space, guided by text embeddings.
Check yourself
- What does the forward diffusion process do?
Show answer
Gradually adds noise to an image — Noise is added step by step until the image is destroyed.
- What is the network trained to predict?
Show answer
The noise added at a given step — Predicting noise lets it denoise.
- Why does latent diffusion use a VAE?
Show answer
To work in a smaller compressed space, which is much cheaper — Denoising small latents is far faster than full pixels.
Go deeper
- Diffusion Models: Generating by Learning to Denoise · The AI Lecture Hall
- Latent Diffusion and Stable Diffusion: Text-to-Image at Scale · The AI Lecture Hall
- Text-to-Image and Text-to-Video Generation: Systems, Control and Provenance · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/diffusion-models-in-depth.html