AI in Motion

Deep LearningIntermediate1:28 video6 chapters

Vanishing Gradients and Residual Connections — lecture notes

Why very deep networks used to be untrainable — gradients shrinking layer by layer — and how skip connections fixed it.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Vanishing Gradients and Residual Connections

For years, adding more layers made networks worse, not better. The culprit was the vanishing gradient. Let us watch it happen, and see the fix.

0:112. Watching gradients vanish

Watching gradients vanish — Vanishing Gradients and Residual Connections

During backpropagation, the gradient is multiplied by each layer’s local slope on its way back. The sigmoid’s slope is at most 0.25. Multiply by 0.25 again and again, and by the first layer only about 0.00006 of the signal remains. Early layers barely learn.

0:303. The maths

The maths — Vanishing Gradients and Residual Connections

With seven layers of sigmoid slopes, the best case is 0.25 to the power 7, which is about 0.00006. Deeper networks make it even smaller. That is why deep sigmoid networks trained so slowly.

0:454. Residual connections

Residual connections — Vanishing Gradients and Residual Connections

Residual networks add skip connections: each block adds its input to its output. The gradient gets a direct highway back through the network. In this illustration it stays around 0.6 even at the first layer. This idea, from ResNet in 2015, made networks with over a hundred layers trainable.

1:065. Other fixes

Other fixes — Vanishing Gradients and Residual Connections

Several tools work together. ReLU activations, careful weight initialisation, normalisation layers, and residual connections. Modern networks, including Transformers, use all of them.

1:166. Recap

Recap — Vanishing Gradients and Residual Connections

To recap. Gradients are multiplied layer by layer. Small slopes make them vanish. Skip connections provide a direct path, and that is how very deep networks became trainable.

Key takeaways

  • Backpropagated gradients are multiplied by each layer’s local derivative.
  • Sigmoid slopes are at most 0.25, so gradients can shrink exponentially (0.25⁷ ≈ 0.00006).
  • Residual (skip) connections give gradients a direct path back.
  • ReLU, good initialisation and normalisation also help.

Check yourself

  1. What is the maximum slope of the sigmoid function?
    Show answer

    0.25 — σ′(z) = σ(z)(1 − σ(z)) peaks at 0.25.

  2. What does a residual connection do?
    Show answer

    Adds a block’s input to its output — y = F(x) + x creates a shortcut for signals and gradients.

  3. Which architecture popularised skip connections in 2015?
    Show answer

    ResNet — ResNet made networks with 100+ layers trainable.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/vanishing-gradients-and-resnets.html