AI in Motion

Deep LearningIntermediate1:26 video6 chapters

Optimisers: SGD, Momentum and Adam — lecture notes

Compare optimisers racing across the same loss surface: noisy SGD, plain gradient descent, momentum and Adam.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Optimisers: SGD, Momentum and Adam

Gradients tell us which way is downhill. Optimisers decide how to step. The choice can make training faster, smoother and more reliable.

0:102. SGD

SGD — Optimisers: SGD, Momentum and Adam

Plain gradient descent uses the whole dataset for every step. Stochastic gradient descent, or SGD, uses a small random mini batch instead. Each step is noisier, as you can see from the wobbly amber path, but much cheaper, so we can take many more steps.

0:293. Momentum

Momentum — Optimisers: SGD, Momentum and Adam

Momentum works like a ball rolling downhill. It keeps a running velocity that averages recent gradients. Zigzags across a valley cancel out, while consistent directions build up speed.

0:424. The race

The race — Optimisers: SGD, Momentum and Adam

Here they race on a narrow valley. Plain gradient descent zigzags. Adam adapts its step size for each direction and heads almost straight for the minimum. Momentum accelerates along the valley and ends up closest by the last step.

0:585. Adam

Adam — Optimisers: SGD, Momentum and Adam

Adam combines momentum with a running average of squared gradients. Dividing by it gives every parameter its own step size. It works well out of the box, so Adam, or its variant AdamW, is a common default.

1:156. Recap

Recap — Optimisers: SGD, Momentum and Adam

To recap. SGD takes cheap noisy steps. Momentum smooths the path. Adam adapts step sizes per parameter. And learning rate schedules help every optimiser.

Key takeaways

  • SGD estimates the gradient on mini-batches: cheap but noisy.
  • Momentum averages past gradients to damp zigzags and speed up progress.
  • Adam adapts the step size for each parameter.
  • AdamW is a common default; learning-rate schedules help all optimisers.

Check yourself

  1. Why is SGD’s path noisy?
    Show answer

    Each step uses a small random mini-batch — Mini-batch gradients are noisy estimates of the full gradient.

  2. What does momentum remember?
    Show answer

    A running average of recent gradients — The velocity term accumulates past gradients.

  3. What makes Adam “adaptive”?
    Show answer

    Each parameter gets its own effective step size — It scales updates by running estimates of squared gradients.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/optimizers-sgd-momentum-adam.html