AI in Motion

Generative AIIntermediate1:29 video6 chapters

Variational Autoencoders — lecture notes

Encode inputs as small probability clouds in a smooth latent space — then walk through that space to generate and morph new data.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Variational Autoencoders

An ordinary autoencoder compresses and rebuilds data. A variational autoencoder, or VAE, organises its compressed codes so neatly that we can generate brand new data from them.

0:122. Encode and decode

Encode and decode — Variational Autoencoders

Recall the autoencoder: an encoder squeezes the input into a small code, and a decoder rebuilds it. A VAE does the same, but the encoder outputs a small probability cloud, a mean and a spread, instead of one exact point.

0:293. Two forces

Two forces — Variational Autoencoders

Training balances two forces. The reconstruction term says: rebuild the input accurately. The KL term says: keep every code cloud close to a simple standard normal shape. Together they make the latent space smooth and gap-free.

0:454. Walking the latent space

Walking the latent space — Variational Autoencoders

Because the space is smooth, nearby codes decode to similar outputs. In this illustration, one direction changes the shape from circle to square, and the other changes the colour. Walking through the space morphs smoothly between them. Sampling a random point generates something new.

1:045. VAEs today

VAEs today — Variational Autoencoders

VAEs train stably and give meaningful latent spaces, though their samples can look blurry. Today a VAE sits inside Stable Diffusion, compressing images so diffusion can run in a small latent space.

1:186. Recap

Recap — Variational Autoencoders

To recap. VAEs encode inputs as probability clouds, balance reconstruction against a tidy latent space, and let us sample and morph. They also power latent diffusion.

Key takeaways

  • A VAE encoder outputs a distribution (mean and variance) for each input.
  • The loss combines reconstruction error with a KL-divergence regulariser.
  • The resulting latent space is smooth, so sampling and interpolation work.
  • VAEs compress images in latent diffusion models such as Stable Diffusion.

Check yourself

  1. What does a VAE encoder output?
    Show answer

    A mean and a spread (a distribution) — Codes are probability clouds.

  2. What does the KL term encourage?
    Show answer

    Codes close to a standard normal distribution — It keeps the latent space tidy.

  3. Where is a VAE used in Stable Diffusion?
    Show answer

    To compress images into a latent space — Diffusion runs in the VAE’s latent space.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/variational-autoencoders.html