AI in Motion

Computer VisionIntermediate1:35 video6 chapters

Convolutional Neural Networks — lecture notes

Stacks of learned filters, activations and pooling turn pixels into probabilities. Follow data through a CNN and see what each layer learns.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Convolutional Neural Networks

Convolutional neural networks, or CNNs, transformed computer vision. They stack learned filters to go from raw pixels to confident predictions.

0:092. Inside a CNN

Inside a CNN — Convolutional Neural Networks

Follow the data. A 32 by 32 colour image enters. A convolution layer applies 16 learned filters, producing 16 feature maps. Pooling halves the size. Another convolution makes 32 maps, and pooling shrinks again. The spatial size shrinks while the number of channels grows. Finally the maps are flattened, a dense layer combines them, and softmax gives probabilities.

0:333. The feature hierarchy

The feature hierarchy — Convolutional Neural Networks

What does each layer learn? The first layers detect edges and colours. The next combine them into textures and corners. Deeper layers respond to parts like eyes and wheels, and the deepest to whole objects. Nobody programs these features. They emerge during training.

0:514. Why CNNs work

Why CNNs work — Convolutional Neural Networks

Convolution suits images for three reasons. Each filter looks at a small local patch. The same filter is reused across the whole image, so there are far fewer parameters. And a pattern is detected wherever it appears.

1:085. Making a prediction

Making a prediction — Convolutional Neural Networks

At the end, the network outputs a score for every class, and softmax turns the scores into probabilities that add up to 100 percent. Here it is 86 percent confident that this image shows a house.

1:236. Recap

Recap — Convolutional Neural Networks

To recap. Convolution layers learn filters, pooling shrinks the maps, channels grow while size shrinks, and the network builds a hierarchy from edges to whole objects.

Key takeaways

  • CNNs stack convolution, activation and pooling layers, then classify.
  • Spatial size shrinks while the number of feature channels grows.
  • Early layers learn edges; deeper layers learn parts and objects.
  • Local connections and weight sharing make CNNs efficient for images.

Check yourself

  1. What does weight sharing mean in a CNN?
    Show answer

    The same filter is applied across the whole image — One filter scans every position.

  2. As data moves deeper into a typical CNN…
    Show answer

    Spatial size shrinks and channels grow — Pooling/striding reduce size while more filters add channels.

  3. Deeper CNN layers tend to detect…
    Show answer

    More complex parts and objects — Features become more abstract with depth.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/convolutional-neural-networks.html