AI in Motion

Computer VisionIntermediate1:31 video6 chapters

Landmark CNNs: From LeNet to ResNet — lecture notes

The architectures that defined deep vision — and how ImageNet top-5 error fell from 28% to under 4% in five years.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Landmark CNNs: From LeNet to ResNet

A handful of famous networks shaped modern computer vision. Each one introduced an idea we still use today.

0:082. The timeline

The timeline — Landmark CNNs: From LeNet to ResNet

LeNet read handwritten digits in 1998. AlexNet in 2012 showed deep CNNs trained on GPUs could win big. VGG and GoogLeNet went deeper in 2014. ResNet reached 152 layers in 2015. MobileNet brought vision to phones, EfficientNet balanced scaling, and in 2020 the Vision Transformer arrived.

0:283. The ImageNet race

The ImageNet race — Landmark CNNs: From LeNet to ResNet

The ImageNet challenge asked models to recognise a thousand categories. The winning top five error was 28 percent in 2010. AlexNet cut it to 16.4 percent in 2012. By 2015, ResNet reached about 3.6 percent, below one published estimate of human error, around 5 percent.

0:474. Key ideas

Key ideas — Landmark CNNs: From LeNet to ResNet

Each network added an idea. AlexNet used ReLU, dropout and GPUs. VGG stacked small three by three filters. GoogLeNet mixed several filter sizes in parallel. ResNet added skip connections. And MobileNet and EfficientNet optimised accuracy per unit of computation.

1:045. ResNet’s trick

ResNet’s trick — Landmark CNNs: From LeNet to ResNet

ResNet’s trick was the skip connection. Gradients flow back through shortcut paths instead of fading layer by layer, which made networks with over a hundred layers trainable.

1:166. Recap

Recap — Landmark CNNs: From LeNet to ResNet

To recap. LeNet and AlexNet started it. VGG, GoogLeNet and ResNet went deeper and smarter. ImageNet error fell from 28 percent to about 3.6. And efficient networks and transformers are the next chapters.

Key takeaways

  • AlexNet’s 2012 ImageNet win (16.4% top-5 error) launched the deep learning era in vision.
  • VGG, GoogLeNet and ResNet pushed depth and design further.
  • ResNet reached about 3.6% top-5 error in 2015 using skip connections.
  • MobileNet and EfficientNet focus on accuracy per unit of compute.

Check yourself

  1. Which network won ImageNet in 2012 and started the deep learning boom?
    Show answer

    AlexNet — AlexNet cut top-5 error to 16.4%.

  2. What key idea did ResNet introduce?
    Show answer

    Skip (residual) connections — Residual connections let very deep networks train.

  3. What do MobileNet and EfficientNet prioritise?
    Show answer

    Accuracy per unit of computation — They target efficient models, e.g. for phones.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/landmark-cnn-architectures.html