Optimisers: SGD, Momentum and Adam — lecture notes
Compare optimisers racing across the same loss surface: noisy SGD, plain gradient descent, momentum and Adam.
0:001. Introduction

Gradients tell us which way is downhill. Optimisers decide how to step. The choice can make training faster, smoother and more reliable.
0:102. SGD

Plain gradient descent uses the whole dataset for every step. Stochastic gradient descent, or SGD, uses a small random mini batch instead. Each step is noisier, as you can see from the wobbly amber path, but much cheaper, so we can take many more steps.
0:293. Momentum

Momentum works like a ball rolling downhill. It keeps a running velocity that averages recent gradients. Zigzags across a valley cancel out, while consistent directions build up speed.
0:424. The race

Here they race on a narrow valley. Plain gradient descent zigzags. Adam adapts its step size for each direction and heads almost straight for the minimum. Momentum accelerates along the valley and ends up closest by the last step.
0:585. Adam

Adam combines momentum with a running average of squared gradients. Dividing by it gives every parameter its own step size. It works well out of the box, so Adam, or its variant AdamW, is a common default.
1:156. Recap

To recap. SGD takes cheap noisy steps. Momentum smooths the path. Adam adapts step sizes per parameter. And learning rate schedules help every optimiser.
Key takeaways
- SGD estimates the gradient on mini-batches: cheap but noisy.
- Momentum averages past gradients to damp zigzags and speed up progress.
- Adam adapts the step size for each parameter.
- AdamW is a common default; learning-rate schedules help all optimisers.
Check yourself
- Why is SGD’s path noisy?
Show answer
Each step uses a small random mini-batch — Mini-batch gradients are noisy estimates of the full gradient.
- What does momentum remember?
Show answer
A running average of recent gradients — The velocity term accumulates past gradients.
- What makes Adam “adaptive”?
Show answer
Each parameter gets its own effective step size — It scales updates by running estimates of squared gradients.
Go deeper
- Optimisers I: SGD, Momentum and Nesterov Acceleration · The AI Lecture Hall
- Optimisers II: AdaGrad, RMSProp, Adam and AdamW · The AI Lecture Hall
- Learning Rate Schedules, Warm-up and the LR Range Test · The AI Lecture Hall
- Loss Landscapes, Saddle Points and Flat Minima · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/optimizers-sgd-momentum-adam.html