AI in Motion

Deep LearningIntermediate1:23 video6 chapters

Dropout and Regularisation — lecture notes

Big networks memorise. Dropout, weight decay, early stopping and augmentation keep them honest — see dropout flicker neurons on and off.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Dropout and Regularisation

Neural networks can have millions of parameters, enough to memorise their training data. Regularisation techniques stop that and help them generalise.

0:092. Spotting overfitting

Spotting overfitting — Dropout and Regularisation

Here is the warning sign. Training loss keeps falling, but validation loss starts to rise. The network is memorising details of the training examples. Early stopping would halt training at the dashed line.

0:243. Dropout

Dropout — Dropout and Regularisation

Dropout randomly switches off, here, about half of the hidden neurons at every training step. A different random set each time. No neuron can rely on a particular partner, so the network learns more robust, spread-out features. At test time, every neuron is used, with outputs scaled to match.

0:454. Toolkit

Toolkit — Dropout and Regularisation

Other tools include weight decay, which penalises large weights; early stopping; data augmentation, which creates varied training examples; and, most reliable of all, more real data.

0:565. Train vs test

Train vs test — Dropout and Regularisation

During training, dropout is like training many smaller networks that share weights. At test time, using all neurons is like averaging all of them, which usually improves accuracy on new data.

1:106. Recap

Recap — Dropout and Regularisation

To recap. Watch for a widening gap between training and validation loss. Dropout, weight decay, early stopping and augmentation all fight overfitting, and more data beats them all.

Key takeaways

  • Overfitting appears as validation loss rising while training loss keeps falling.
  • Dropout randomly disables neurons during training so features become robust.
  • At test time all neurons are used, with outputs scaled.
  • Weight decay, early stopping, augmentation and more data also reduce overfitting.

Check yourself

  1. What does dropout do during training?
    Show answer

    Randomly switches off some neurons at each step — A new random subset is dropped every step.

  2. Is dropout applied when making predictions on new data?
    Show answer

    No — all neurons are used, with scaled outputs — Dropout is a training-time technique.

  3. What does early stopping use to decide when to stop?
    Show answer

    Validation loss — Training stops when validation performance stops improving.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/dropout-and-regularisation.html