Convolutional Neural Networks
Stacks of learned filters, activations and pooling turn pixels into probabilities. Follow data through a CNN and see what each layer learns.
π Illustrated notes Β· every chapter as a picture Β· printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. Convolutional neural networks, or CNNs, transformed computer vision. They stack learned filters to go from raw pixels to confident predictions.
Inside a CNN. Follow the data. A 32 by 32 colour image enters. A convolution layer applies 16 learned filters, producing 16 feature maps. Pooling halves the size. Another convolution makes 32 maps, and pooling shrinks again. The spatial size shrinks while the number of channels grows. Finally the maps are flattened, a dense layer combines them, and softmax gives probabilities.
The feature hierarchy. What does each layer learn? The first layers detect edges and colours. The next combine them into textures and corners. Deeper layers respond to parts like eyes and wheels, and the deepest to whole objects. Nobody programs these features. They emerge during training.
Why CNNs work. Convolution suits images for three reasons. Each filter looks at a small local patch. The same filter is reused across the whole image, so there are far fewer parameters. And a pattern is detected wherever it appears.
Making a prediction. At the end, the network outputs a score for every class, and softmax turns the scores into probabilities that add up to 100 percent. Here it is 86 percent confident that this image shows a house.
Recap. To recap. Convolution layers learn filters, pooling shrinks the maps, channels grow while size shrinks, and the network builds a hierarchy from edges to whole objects.