Computer Vision
Pixels, convolution, edges, CNNs, classic architectures, augmentation, detection, segmentation, ViTs and pose.
15 animated lectures · 41 minutes
▶ Start with lecture 1How Computers See Images
To a computer, a picture is a grid of numbers. Zoom into the pixels, read their values and split a colour image into red, green and blue channels.
Convolution and Image Filters
Slide a small grid of numbers over an image, multiply and add. See edge-detection, blur and sharpen filters computed cell by cell.
Edge Detection with Sobel Filters
Edges are where brightness changes quickly. Compute horizontal and vertical gradients with Sobel filters and combine them into an edge map.
Convolutional Neural Networks
Stacks of learned filters, activations and pooling turn pixels into probabilities. Follow data through a CNN and see what each layer learns.
Pooling, Stride and Padding
How CNNs shrink feature maps: max pooling, average pooling, stride and padding — with the numbers computed in front of you.
Landmark CNNs: From LeNet to ResNet
The architectures that defined deep vision — and how ImageNet top-5 error fell from 28% to under 4% in five years.
Image Classification End to End
From a labelled dataset to a trained classifier: the full pipeline, softmax probabilities and how to judge the results.
Data Augmentation for Vision
One image becomes many training examples: flips, rotations, crops, lighting changes and noise teach models what really matters.
Transfer Learning for Vision
Reuse a network trained on millions of images: freeze its layers, add a new head, and fine-tune with only a small dataset of your own.
Object Detection: Boxes, IoU and NMS
Find every object and draw a box around it. From sliding windows to YOLO-style grids, IoU and non-maximum suppression.
Image Segmentation: Every Pixel Labelled
Semantic segmentation labels every pixel by class; instance segmentation separates each object. Plus the U-Net architecture that made it practical.
Vision Transformers: Images as Patches
Cut an image into patches, treat them like words and let a Transformer attend between them. How ViTs work and when they beat CNNs.
Human Pose Estimation
Find a person’s joints — shoulders, elbows, knees — and connect them into a skeleton that can be tracked over time.
Convolutional Neural Networks: A Deep Dive
From pixels to predictions: why convolution, kernels and feature maps, stride, padding and channels, parameter counts, pooling, receptive fields, landmark architectures, residual connections, augmentation and transfer learning.
Object Detection and Segmentation: A Deep Dive
Finding and outlining objects: sliding windows, R-CNN to Faster R-CNN, YOLO and one-stage detectors, anchors, IoU, non-maximum suppression, mAP, focal loss, DETR, semantic, instance and panoptic segmentation, U-Net and Segment Anything.