110 animated lectures · 36 deep dives · 502 minutes
See how AI actually works.
Animated video lectures on AI, Machine Learning, Deep Learning, Computer Vision, NLP, Generative AI, LLMs, Reinforcement Learning, MLOps and the Mathematics of ML — every lesson is an animated video with narration, subtitles, chapters and a quiz. Short lessons explain one idea in minutes; deep dives of ten minutes or more take you from intuition to the maths and the code.
- 🎙️ Voice narration
- 💬 Subtitles & transcripts
- 🧩 Chapters & quizzes
- 📄 Printable illustrated notes
- 📚 Linked to The AI Lecture Hall
Live preview — every frame is drawn in real time in your browser
Artificial Intelligence
Agents, search, games, optimisation, probability, language models, attention and image generation.
14 lectures · 41 min →◈Machine Learning
Regression, gradient descent, classifiers, trees, SVMs, clustering, PCA, overfitting and evaluation.
14 lectures · 39 min →⬡Deep Learning
Neurons, networks, activations, training, backpropagation, optimisers, RNNs, embeddings, autoencoders and GANs.
14 lectures · 40 min →◐Computer Vision
Pixels, convolution, edges, CNNs, classic architectures, augmentation, detection, segmentation, ViTs and pose.
15 lectures · 41 min →❝Natural Language Processing
Tokenization, TF-IDF, n-grams, word2vec, classification, NER, translation, positional encoding, BERT vs GPT, semantic search and speech.
14 lectures · 38 min →✦Generative AI
VAEs, diffusion, guidance, LLM training, decoding, prompting, RAG, LoRA, quantisation, mixture of experts, agents and multimodal models.
15 lectures · 40 min →❖Large Language Models
Deep dives into transformer internals, attention maths, pre-training at scale, alignment (SFT, RLHF, DPO), inference engineering and building LLM applications.
6 lectures · 66 min →♞Reinforcement Learning
MDPs and returns, value and policy iteration, Monte Carlo and TD learning, Q-learning vs SARSA, bandits, deep Q-networks, policy gradients and PPO.
6 lectures · 66 min →⚙MLOps & Engineering
The ML lifecycle, data and experiment management, Docker and model serving, CI/CD and deployment strategies, monitoring and drift, A/B testing and responsible operations.
6 lectures · 64 min →∑Mathematics for ML
Vectors and dot products, matrices as transformations, eigenvectors, SVD and PCA, calculus and gradients, probability, likelihood and information theory.
6 lectures · 66 min →All animated lectures
What Is Artificial Intelligence?
A clear, visual introduction: what AI is, how modern AI learns from data, and how AI, machine learning and deep learning fit together.
A Short History of AI
From Alan Turing to ChatGPT in one animated timeline — the breakthroughs, the “AI winters” and the ideas that changed everything.
Intelligent Agents: Perceive, Decide, Act
Watch a robot vacuum perceive its world, decide and act — the agent loop that underlies everything from thermostats to AI assistants.
Breadth-First vs Depth-First Search
Watch two classic search algorithms explore the same maze — one in ripples, one in deep dives — and see why only one guarantees the shortest path.
A* Search: Smarter Pathfinding
A* combines the cost so far with an estimate of the cost remaining. Watch it head for the goal and find the same shortest path while exploring fewer cells.
Minimax and Alpha–Beta Pruning
How game-playing AI thinks ahead: MAX and MIN take turns, values flow up the tree, and alpha–beta pruning skips branches that cannot matter.
Hill Climbing and Simulated Annealing
Local search climbs towards better solutions — but gets stuck on local peaks. Simulated annealing adds controlled randomness to escape.
Genetic Algorithms: Evolution in Code
Selection, crossover and mutation: watch a population of bit strings evolve towards a perfect solution, generation by generation.
Bayes’ Theorem: Reasoning Under Uncertainty
A positive medical test that is “90% accurate” — so why is the chance of being sick only about 32%? Bayes’ theorem explained with 200 people.
How Large Language Models Work
ChatGPT, Claude and friends predict the next token over and over. Watch a model choose words from probabilities, one token at a time.
Attention and Transformers
Self-attention lets every word look at every other word. See how “it” finds “animal”, and how attention matrices power the Transformer.
How AI Image Generators Work
Diffusion models turn pure noise into a picture by removing a little noise at a time, guided by your text prompt.
Search and Problem Solving: A Deep Dive
How AI systems find solutions: state spaces, breadth-first, depth-first, uniform-cost, greedy and A* search, admissible heuristics, local search, simulated annealing, genetic algorithms and constraint satisfaction.
Game-Playing AI: From Minimax to AlphaZero
How machines learned to beat world champions: game trees, minimax, evaluation functions, alpha–beta pruning, Deep Blue, Monte Carlo tree search, UCT, and AlphaGo and AlphaZero’s marriage of search and learning.
What Is Machine Learning?
Instead of writing rules, show the computer examples. The core idea of machine learning, its three main types and the standard workflow.
Linear Regression, Visually
Fit a straight line to data with gradient descent. Watch the residuals shrink and the mean squared error fall from 9.45 to 0.55.
Gradient Descent and the Learning Rate
How almost every ML model learns: follow the slope downhill. See small, good and too-large learning rates, and how optimisers like Adam move on a loss surface.
Linear Classifiers: Logistic Regression and the Perceptron
Draw a line that separates two classes. Watch logistic regression learn a probabilistic boundary and the perceptron nudge its line towards mistakes.
k-Nearest Neighbours: Learning by Similarity
Classify a new point by asking its closest neighbours to vote. Simple, intuitive and a great first classifier.
Decision Trees and Random Forests
A tree learns yes/no questions that carve the data into pure regions. A forest of trees votes for more robust predictions.
Support Vector Machines and the Kernel Trick
Of all the lines that separate two classes, pick the widest street. Then lift the data into a new dimension to separate what a line cannot.
k-Means Clustering, Step by Step
Unsupervised learning in action: assign points to the nearest centroid, move each centroid to its points’ average, repeat.
Principal Component Analysis (PCA)
Find the direction in which data varies most, then project onto it. Watch two dimensions become one while keeping 93.5% of the variance.
Overfitting, Underfitting and Bias–Variance
Too simple, just right, too complex: see three models on the same data and why validation error — not training error — tells the truth.
Train/Test Splits and Cross-Validation
Why data is split into training, validation and test sets — and how k-fold cross-validation gives a more reliable score.
Confusion Matrix, Precision, Recall and ROC
Accuracy alone can mislead. Build a confusion matrix, compute precision, recall and F1, then move the threshold and trace an ROC curve.
Gradient Descent and Optimisation: A Deep Dive
Everything about how models are fitted: loss functions, learning rates, batch vs stochastic gradient descent, momentum, Adam and AdamW, schedules, conditioning and feature scaling, saddle points and practical tuning.
Evaluating Machine Learning Models: A Deep Dive
How to know whether a model is really good: train/validation/test splits, cross-validation, confusion matrices, precision, recall and F1, ROC and PR curves, calibration, regression metrics, bias–variance, learning curves and leakage.
Artificial Neurons and the Perceptron
Inside a single artificial neuron: weighted inputs, a sum, a bias and an activation — plus the perceptron that learns from its mistakes.
Neural Networks: The Forward Pass
Stack neurons into layers and connect them. Watch signals flow from input to output and become class probabilities.
Activation Functions: Sigmoid, Tanh and ReLU
Why networks need non-linearity, and how sigmoid, tanh, ReLU, Leaky ReLU and GELU differ — drawn live.
Loss Functions and the Training Loop
How a network measures its mistakes and improves: mini-batches, forward pass, loss, backward pass and update — repeated thousands of times.
Backpropagation, Step by Step
The chain rule in action: compute values forward, then pass gradients backward through a tiny computational graph — with real numbers.
Optimisers: SGD, Momentum and Adam
Compare optimisers racing across the same loss surface: noisy SGD, plain gradient descent, momentum and Adam.
Vanishing Gradients and Residual Connections
Why very deep networks used to be untrainable — gradients shrinking layer by layer — and how skip connections fixed it.
Dropout and Regularisation
Big networks memorise. Dropout, weight decay, early stopping and augmentation keep them honest — see dropout flicker neurons on and off.
Recurrent Networks and LSTMs
Networks with memory: an RNN reads a sentence word by word, carrying a hidden state; an LSTM adds gates to remember for longer.
Embeddings: Meaning as Vectors
Words become points in space where similar meanings sit close together — and directions carry meaning, as in king − man + woman ≈ queen.
Autoencoders: Compress and Reconstruct
Squeeze an image through a tiny bottleneck and rebuild it. Autoencoders learn compact representations — and can clean up noisy inputs.
Generative Adversarial Networks (GANs)
A forger and a detective train together. Watch the generator’s fake distribution move until it matches real data.
Training Deep Neural Networks: A Practical Deep Dive
The full recipe for training a network well: forward pass, softmax and cross-entropy, backpropagation, initialisation, vanishing gradients and skip connections, normalisation, optimisers, regularisation and a debugging checklist.
From RNNs to Transformers: Sequence Models in Depth
How neural networks learned to handle sequences: recurrent networks and backpropagation through time, vanishing gradients, LSTM and GRU gates, encoder–decoder translation, attention, and why transformers replaced recurrence.
How Computers See Images
To a computer, a picture is a grid of numbers. Zoom into the pixels, read their values and split a colour image into red, green and blue channels.
Convolution and Image Filters
Slide a small grid of numbers over an image, multiply and add. See edge-detection, blur and sharpen filters computed cell by cell.
Edge Detection with Sobel Filters
Edges are where brightness changes quickly. Compute horizontal and vertical gradients with Sobel filters and combine them into an edge map.
Convolutional Neural Networks
Stacks of learned filters, activations and pooling turn pixels into probabilities. Follow data through a CNN and see what each layer learns.
Pooling, Stride and Padding
How CNNs shrink feature maps: max pooling, average pooling, stride and padding — with the numbers computed in front of you.
Landmark CNNs: From LeNet to ResNet
The architectures that defined deep vision — and how ImageNet top-5 error fell from 28% to under 4% in five years.
Image Classification End to End
From a labelled dataset to a trained classifier: the full pipeline, softmax probabilities and how to judge the results.
Data Augmentation for Vision
One image becomes many training examples: flips, rotations, crops, lighting changes and noise teach models what really matters.
Transfer Learning for Vision
Reuse a network trained on millions of images: freeze its layers, add a new head, and fine-tune with only a small dataset of your own.
Object Detection: Boxes, IoU and NMS
Find every object and draw a box around it. From sliding windows to YOLO-style grids, IoU and non-maximum suppression.
Image Segmentation: Every Pixel Labelled
Semantic segmentation labels every pixel by class; instance segmentation separates each object. Plus the U-Net architecture that made it practical.
Vision Transformers: Images as Patches
Cut an image into patches, treat them like words and let a Transformer attend between them. How ViTs work and when they beat CNNs.
Human Pose Estimation
Find a person’s joints — shoulders, elbows, knees — and connect them into a skeleton that can be tracked over time.
Convolutional Neural Networks: A Deep Dive
From pixels to predictions: why convolution, kernels and feature maps, stride, padding and channels, parameter counts, pooling, receptive fields, landmark architectures, residual connections, augmentation and transfer learning.
Object Detection and Segmentation: A Deep Dive
Finding and outlining objects: sliding windows, R-CNN to Faster R-CNN, YOLO and one-stage detectors, anchors, IoU, non-maximum suppression, mAP, focal loss, DETR, semantic, instance and panoptic segmentation, U-Net and Segment Anything.
What Is Natural Language Processing?
How computers read, understand and generate human language — the tasks, the pipeline and how the field moved from rules to neural networks.
Tokenization and Byte-Pair Encoding
Words, characters or sub-words? See three ways to tokenise text, then watch byte-pair encoding learn sub-words from a tiny corpus.
Bag of Words and TF-IDF
The classic way to turn documents into numbers: count words, then weigh them by how rare they are. Computed live on three tiny documents.
N-gram Language Models
Predict the next word by counting word pairs. A bigram model built from a tiny corpus — the ancestor of today’s LLMs.
Word2vec and Word Embeddings
You shall know a word by the company it keeps. See how skip-gram turns context windows into training pairs — and meaning into vectors.
Sentiment Analysis and Text Classification
Is this review positive or negative? See how word evidence adds up, why negation is tricky, and how modern classifiers are built.
Named Entity Recognition
Find people, organisations, places and dates in text and tag every token with BIO labels.
Sequence-to-Sequence Models and Translation
An encoder reads a sentence, a decoder writes the translation — and attention lets it look back at exactly the right words.
Positional Encoding: Teaching Transformers Word Order
Attention ignores order, so Transformers add position information. See the sine-and-cosine pattern that gives every position a unique fingerprint.
BERT and GPT: Two Kinds of Language Models
BERT reads in both directions to understand; GPT reads left to right to generate. See their training games and attention masks.
Semantic Search with Embeddings
Search by meaning, not matching words. Embed documents and queries, then find the nearest neighbours.
Speech Recognition: From Sound to Text
How a voice becomes words: waveforms, spectrograms and neural models that turn frequency patterns into text.
Word Embeddings: A Deep Dive
How words became vectors: one-hot and TF-IDF, the distributional hypothesis, word2vec skip-gram with negative sampling, analogies, GloVe and fastText, contextual embeddings from BERT, sentence embeddings for search, and bias.
Text Classification from Start to Finish
Build a real text classifier: framing and labelling, tokenisation and TF-IDF features, Naive Bayes and logistic-regression baselines, cross-validation and metrics, imbalance, fine-tuning transformers, zero-shot LLMs, error analysis and monitoring.
Generative vs Discriminative Models
One kind of model learns where the boundary is; the other learns what the data looks like — and can create new examples.
Variational Autoencoders
Encode inputs as small probability clouds in a smooth latent space — then walk through that space to generate and morph new data.
Diffusion Models in Depth
The forward process adds noise on a schedule; a neural network learns to reverse it. See the real DDPM noise schedule and latent diffusion.
Classifier-Free Guidance
How image generators follow prompts more closely: combine a conditional and an unconditional prediction and push further in the prompt’s direction.
How Large Language Models Are Trained
Pre-training on vast text, instruction tuning, and learning from human preferences — plus the scaling laws that made models grow.
Decoding: Temperature, Top-k and Top-p
How a language model picks each word from its probabilities — and how temperature, top-k and top-p change its personality.
Prompt Engineering and Chain-of-Thought
Clear roles, context, examples and step-by-step reasoning: how to get much better answers from language models.
Retrieval-Augmented Generation (RAG)
Give a language model the right documents at the right moment: retrieve relevant passages, then generate a grounded answer with citations.
LoRA and Parameter-Efficient Fine-Tuning
Fine-tune a huge model by training two tiny matrices. See why LoRA needs well under 1% of the parameters of full fine-tuning.
Quantisation: Smaller, Faster Models
Store each weight in fewer bits. See weights snap to int8 and int4 levels, and how a 7B model shrinks from 28 GB to 3.5 GB.
Mixture of Experts
A router sends each token to a few specialised experts, so a model can have many parameters while using only a fraction per token.
AI Agents and Tool Use
Language models that act: plan, call tools like search or calculators, read the results and continue until the task is done.
Multimodal Models and CLIP
Put images and text in one shared space by pulling matching pairs together — the idea behind CLIP, image search and vision-language assistants.
Diffusion Models: The Complete Deep Dive
How image generators really work: the forward noising process and its schedule, the noise-prediction objective, U-Net denoisers, DDPM vs DDIM sampling, latent diffusion with a VAE, text conditioning with CLIP and cross-attention, classifier-free guidance, ControlNet and fast samplers.
GANs and VAEs: A Deep Dive into Generative Models
Two classic ways to generate data: autoencoders and variational autoencoders (ELBO, KL, reparameterisation), and generative adversarial networks (the minimax game, mode collapse, DCGAN, WGAN, conditional and cycle GANs, StyleGAN) — with evaluation and a comparison with diffusion.
Inside a Large Language Model
Follow a sentence through a GPT-style model: tokens, embeddings, the residual stream, attention and MLP blocks, layer norms, unembedding and sampling — with real parameter counts.
The Attention Mechanism in Depth
Queries, keys and values; scaled dot-product attention computed by hand; masking; multi-head attention; positional encodings and RoPE; the quadratic cost and how FlashAttention, GQA and sliding windows tame it.
Pre-training LLMs at Scale
The recipe behind base models: web-scale data pipelines, next-token loss and perplexity, scaling laws and compute budgets, Chinchilla, distributed training, learning-rate schedules and what can go wrong.
Aligning LLMs: Instruction Tuning, RLHF and DPO
How a text predictor becomes a helpful assistant: supervised fine-tuning, preference data, reward models, RLHF with PPO and a KL leash, DPO, AI feedback, reward hacking and evaluation.
LLM Inference Engineering
Why serving LLMs is hard and how it is made fast and cheap: prefill vs decode, memory bandwidth, the KV cache, continuous batching, PagedAttention, quantisation, speculative decoding and caching.
Building with LLMs: Prompting, RAG and Agents
The practical toolkit: prompt design and in-context learning, chain-of-thought, structured outputs, retrieval-augmented generation end to end, tool-using agents, evaluation, prompt injection and cost.
Reinforcement Learning Foundations: Agents, MDPs and Returns
How an agent learns from rewards: the agent–environment loop, Markov decision processes, discounted returns, policies and value functions — the vocabulary behind every RL algorithm.
Dynamic Programming: Value Iteration and Policy Iteration
When the rules of the world are known, the Bellman equations can be solved exactly. Watch value iteration spread value from the goal and policy iteration converge in five rounds.
Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning
Model-free reinforcement learning: estimate values from sampled episodes, bootstrap with temporal-difference updates, and compare on-policy SARSA with off-policy Q-learning on the famous cliff.
Exploration and Multi-Armed Bandits
The purest form of the explore–exploit dilemma: greedy, ε-greedy, UCB and Thompson sampling, regret, and where bandits run in the real world — from A/B tests to recommendations.
Deep Q-Networks: Reinforcement Learning Meets Deep Learning
How DQN learned Atari from pixels: function approximation, the deadly triad, experience replay, target networks, and the improvements that became Rainbow.
Policy Gradients, Actor–Critic and PPO
Optimise the policy directly: the policy-gradient theorem and REINFORCE, baselines and advantages, actor–critic methods, PPO’s clipped objective — and how the same ideas align large language models.
What is MLOps? The Machine Learning Lifecycle
Why a good model in a notebook is only the start: the ML lifecycle, hidden technical debt, training–serving skew, MLOps maturity levels, the tool landscape and the principles that keep models working in production.
Data Pipelines, Versioning and Experiment Tracking
Treat data like code: validation, versioning with DVC, leakage-safe splits, feature stores and point-in-time correctness, experiment tracking with MLflow, reproducibility and model registries.
Packaging and Serving Models
From a model file to a reliable service: batch vs online inference, a FastAPI endpoint, Docker images, Kubernetes and autoscaling, latency percentiles and queueing, GPU batching, optimisation and edge deployment.
CI/CD for ML and Safe Deployment Strategies
Automate the path to production: testing code, data and models; CI/CD/CT pipelines with quality gates; champion–challenger evaluation; shadow, canary and blue–green deployments; and fast rollback.
Monitoring, Drift and Retraining
Why models decay and how to catch it: what to monitor, data drift vs concept drift, the population stability index, delayed labels, alerting, incident response and retraining strategies.
A/B Testing and Responsible ML in Production
Measure real impact and operate responsibly: online vs offline evaluation, A/B test design, significance and the peeking trap, sample sizes, model cards, fairness, privacy, explainability and governance.
Vectors, Norms and the Dot Product
A deep dive into the object every model is built from: what vectors are, how to add, scale and measure them, and why the dot product powers neurons, attention and semantic search.
Matrices as Transformations
See matrices as machines that transform space: matrix–vector products, composition, determinants, inverses and rank — and why every neural-network layer is a matrix.
Eigenvectors, SVD and PCA
Find the directions a matrix does not turn, break any matrix into rotate–stretch–rotate, and use it to compress data with principal component analysis.
Calculus for Machine Learning: Derivatives, Gradients and the Chain Rule
How models learn by following slopes: derivatives, partial derivatives, gradients, gradient descent, saddle points and the chain rule that makes backpropagation possible.
Probability and Distributions for ML
Random variables, expectation and variance, the Bernoulli, binomial and normal distributions, the central limit theorem, Monte Carlo methods and Markov chains — with live simulations.
Likelihood, Bayes and Information Theory
Why models minimise cross-entropy: maximum likelihood, priors and MAP, Bayesian updating, entropy, cross-entropy and KL divergence — the statistics hiding inside every loss function.
No lectures match your search.
Watch, check, go deeper
Short videos explain one idea visually in a few minutes; deep dives spend ten minutes or more going from intuition to formulas, worked examples and code. Test yourself with the quiz, then continue with the in-depth written lectures in The AI Lecture Hall.
- Watch the animation with narration — pause, rewind, jump by chapter or change speed.
- Read the transcript and key takeaways under every video.
- Check your understanding with the quiz.
- Go deeper with the linked university-level lectures.