Generative AI
VAEs, diffusion, guidance, LLM training, decoding, prompting, RAG, LoRA, quantisation, mixture of experts, agents and multimodal models.
15 animated lectures · 40 minutes
▶ Start with lecture 1Generative vs Discriminative Models
One kind of model learns where the boundary is; the other learns what the data looks like — and can create new examples.
Variational Autoencoders
Encode inputs as small probability clouds in a smooth latent space — then walk through that space to generate and morph new data.
Diffusion Models in Depth
The forward process adds noise on a schedule; a neural network learns to reverse it. See the real DDPM noise schedule and latent diffusion.
Classifier-Free Guidance
How image generators follow prompts more closely: combine a conditional and an unconditional prediction and push further in the prompt’s direction.
How Large Language Models Are Trained
Pre-training on vast text, instruction tuning, and learning from human preferences — plus the scaling laws that made models grow.
Decoding: Temperature, Top-k and Top-p
How a language model picks each word from its probabilities — and how temperature, top-k and top-p change its personality.
Prompt Engineering and Chain-of-Thought
Clear roles, context, examples and step-by-step reasoning: how to get much better answers from language models.
Retrieval-Augmented Generation (RAG)
Give a language model the right documents at the right moment: retrieve relevant passages, then generate a grounded answer with citations.
LoRA and Parameter-Efficient Fine-Tuning
Fine-tune a huge model by training two tiny matrices. See why LoRA needs well under 1% of the parameters of full fine-tuning.
Quantisation: Smaller, Faster Models
Store each weight in fewer bits. See weights snap to int8 and int4 levels, and how a 7B model shrinks from 28 GB to 3.5 GB.
Mixture of Experts
A router sends each token to a few specialised experts, so a model can have many parameters while using only a fraction per token.
AI Agents and Tool Use
Language models that act: plan, call tools like search or calculators, read the results and continue until the task is done.
Multimodal Models and CLIP
Put images and text in one shared space by pulling matching pairs together — the idea behind CLIP, image search and vision-language assistants.
Diffusion Models: The Complete Deep Dive
How image generators really work: the forward noising process and its schedule, the noise-prediction objective, U-Net denoisers, DDPM vs DDIM sampling, latent diffusion with a VAE, text conditioning with CLIP and cross-attention, classifier-free guidance, ControlNet and fast samplers.
GANs and VAEs: A Deep Dive into Generative Models
Two classic ways to generate data: autoencoders and variational autoencoders (ELBO, KL, reparameterisation), and generative adversarial networks (the minimax game, mode collapse, DCGAN, WGAN, conditional and cycle GANs, StyleGAN) — with evaluation and a comparison with diffusion.