Pooling, Stride and Padding — lecture notes
How CNNs shrink feature maps: max pooling, average pooling, stride and padding — with the numbers computed in front of you.
0:001. Introduction

CNNs gradually shrink their feature maps so later layers see a bigger picture with less computation. Pooling, stride and padding control how.
0:102. Max pooling

Max pooling slides a two by two window with a stride of two, and keeps only the largest value in each window. Six, five, seven and nine. The four by four map becomes two by two: a quarter of the size, keeping the strongest signals.
0:293. Average pooling

Average pooling keeps the mean of each window instead: 3.5, 2, 3.25 and 6.75. It is smoother, and a global version, averaging a whole feature map to one number, is common at the end of modern networks.
0:454. Stride and padding

Stride is how far the filter moves each step. A stride of two roughly halves the output. Padding adds a border of zeros so pixels at the edge are not ignored, and same padding keeps the size unchanged.
1:025. Output size

The output size is n minus k plus two p, divided by the stride, plus one. For a 32 pixel input, a three by three kernel, padding one and stride one, the output stays 32.
1:176. Recap

To recap. Max pooling keeps the strongest value, average pooling the mean. Stride sets the step and padding protects the edges. And one formula predicts the output size.
Key takeaways
- 2 × 2 max pooling with stride 2 keeps the largest value in each window and quarters the map.
- In the demo: max pooling gives 6, 5, 7, 9; average pooling gives 3.5, 2, 3.25, 6.75.
- Stride sets how far a filter moves; padding adds a border of zeros.
- Output size = (n − k + 2p) / s + 1.
Check yourself
- What does max pooling keep from each window?
Show answer
The largest value — Max pooling keeps the strongest activation.
- With n = 32, k = 3, p = 1, s = 1, what is the output size?
Show answer
32 — (32 − 3 + 2) / 1 + 1 = 32.
- Why add padding?
Show answer
So pixels at the image border are not lost — Padding lets filters cover edge pixels and controls output size.
Go deeper
- Padding, Stride, Pooling and Receptive Fields · The AI Lecture Hall
- Convolutional Neural Networks: The Core Ideas · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/pooling-stride-and-padding.html