Variational Autoencoders
Encode inputs as small probability clouds in a smooth latent space — then walk through that space to generate and morph new data.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. An ordinary autoencoder compresses and rebuilds data. A variational autoencoder, or VAE, organises its compressed codes so neatly that we can generate brand new data from them.
Encode and decode. Recall the autoencoder: an encoder squeezes the input into a small code, and a decoder rebuilds it. A VAE does the same, but the encoder outputs a small probability cloud, a mean and a spread, instead of one exact point.
Two forces. Training balances two forces. The reconstruction term says: rebuild the input accurately. The KL term says: keep every code cloud close to a simple standard normal shape. Together they make the latent space smooth and gap-free.
Walking the latent space. Because the space is smooth, nearby codes decode to similar outputs. In this illustration, one direction changes the shape from circle to square, and the other changes the colour. Walking through the space morphs smoothly between them. Sampling a random point generates something new.
VAEs today. VAEs train stably and give meaningful latent spaces, though their samples can look blurry. Today a VAE sits inside Stable Diffusion, compressing images so diffusion can run in a small latent space.
Recap. To recap. VAEs encode inputs as probability clouds, balance reconstruction against a tidy latent space, and let us sample and morph. They also power latent diffusion.