AI in Motion
◐

Computer Vision

Pixels, convolution, edges, CNNs, classic architectures, augmentation, detection, segmentation, ViTs and pose.

15 animated lectures · 41 minutes

▶ Start with lecture 1
1:36
Computer Vision · 01

How Computers See Images

To a computer, a picture is a grid of numbers. Zoom into the pixels, read their values and split a colour image into red, green and blue channels.

Beginner · 6 chapters · 3-question quiz
1:33
Computer Vision · 02

Convolution and Image Filters

Slide a small grid of numbers over an image, multiply and add. See edge-detection, blur and sharpen filters computed cell by cell.

Beginner · 6 chapters · 3-question quiz
1:18
Computer Vision · 03

Edge Detection with Sobel Filters

Edges are where brightness changes quickly. Compute horizontal and vertical gradients with Sobel filters and combine them into an edge map.

Beginner · 5 chapters · 3-question quiz
1:35
Computer Vision · 04

Convolutional Neural Networks

Stacks of learned filters, activations and pooling turn pixels into probabilities. Follow data through a CNN and see what each layer learns.

Intermediate · 6 chapters · 3-question quiz
1:29
Computer Vision · 05

Pooling, Stride and Padding

How CNNs shrink feature maps: max pooling, average pooling, stride and padding — with the numbers computed in front of you.

Beginner · 6 chapters · 3-question quiz
1:31
Computer Vision · 06

Landmark CNNs: From LeNet to ResNet

The architectures that defined deep vision — and how ImageNet top-5 error fell from 28% to under 4% in five years.

Intermediate · 6 chapters · 3-question quiz
1:20
Computer Vision · 07

Image Classification End to End

From a labelled dataset to a trained classifier: the full pipeline, softmax probabilities and how to judge the results.

Beginner · 6 chapters · 3-question quiz
1:24
Computer Vision · 08

Data Augmentation for Vision

One image becomes many training examples: flips, rotations, crops, lighting changes and noise teach models what really matters.

Beginner · 6 chapters · 3-question quiz
1:22
Computer Vision · 09

Transfer Learning for Vision

Reuse a network trained on millions of images: freeze its layers, add a new head, and fine-tune with only a small dataset of your own.

Beginner · 6 chapters · 3-question quiz
1:37
Computer Vision · 10

Object Detection: Boxes, IoU and NMS

Find every object and draw a box around it. From sliding windows to YOLO-style grids, IoU and non-maximum suppression.

Intermediate · 6 chapters · 3-question quiz
1:25
Computer Vision · 11

Image Segmentation: Every Pixel Labelled

Semantic segmentation labels every pixel by class; instance segmentation separates each object. Plus the U-Net architecture that made it practical.

Intermediate · 6 chapters · 3-question quiz
1:31
Computer Vision · 12

Vision Transformers: Images as Patches

Cut an image into patches, treat them like words and let a Transformer attend between them. How ViTs work and when they beat CNNs.

Intermediate · 6 chapters · 3-question quiz
1:08
Computer Vision · 13

Human Pose Estimation

Find a person’s joints — shoulders, elbows, knees — and connect them into a skeleton that can be tracked over time.

Intermediate · 5 chapters · 3-question quiz
Deep dive10:52
Computer Vision · 14

Convolutional Neural Networks: A Deep Dive

From pixels to predictions: why convolution, kernels and feature maps, stride, padding and channels, parameter counts, pooling, receptive fields, landmark architectures, residual connections, augmentation and transfer learning.

Intermediate · 32 chapters · 5-question quiz
Deep dive10:59
Computer Vision · 15

Object Detection and Segmentation: A Deep Dive

Finding and outlining objects: sliding windows, R-CNN to Faster R-CNN, YOLO and one-stage detectors, anchors, IoU, non-maximum suppression, mAP, focal loss, DETR, semantic, instance and panoptic segmentation, U-Net and Segment Anything.

Advanced · 32 chapters · 5-question quiz