AI in Motion

Variational Autoencoders

Generative AIIntermediate1:296 chapters

Encode inputs as small probability clouds in a smooth latent space — then walk through that space to generate and morph new data.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 What does a VAE encoder output?
Q2 What does the KL term encourage?
Q3 Where is a VAE used in Stable Diffusion?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. An ordinary autoencoder compresses and rebuilds data. A variational autoencoder, or VAE, organises its compressed codes so neatly that we can generate brand new data from them.

Encode and decode. Recall the autoencoder: an encoder squeezes the input into a small code, and a decoder rebuilds it. A VAE does the same, but the encoder outputs a small probability cloud, a mean and a spread, instead of one exact point.

Two forces. Training balances two forces. The reconstruction term says: rebuild the input accurately. The KL term says: keep every code cloud close to a simple standard normal shape. Together they make the latent space smooth and gap-free.

Walking the latent space. Because the space is smooth, nearby codes decode to similar outputs. In this illustration, one direction changes the shape from circle to square, and the other changes the colour. Walking through the space morphs smoothly between them. Sampling a random point generates something new.

VAEs today. VAEs train stably and give meaningful latent spaces, though their samples can look blurry. Today a VAE sits inside Stable Diffusion, compressing images so diffusion can run in a small latent space.

Recap. To recap. VAEs encode inputs as probability clouds, balance reconstruction against a tidy latent space, and let us sample and morph. They also power latent diffusion.