Optimisers: SGD, Momentum and Adam
Compare optimisers racing across the same loss surface: noisy SGD, plain gradient descent, momentum and Adam.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. Gradients tell us which way is downhill. Optimisers decide how to step. The choice can make training faster, smoother and more reliable.
SGD. Plain gradient descent uses the whole dataset for every step. Stochastic gradient descent, or SGD, uses a small random mini batch instead. Each step is noisier, as you can see from the wobbly amber path, but much cheaper, so we can take many more steps.
Momentum. Momentum works like a ball rolling downhill. It keeps a running velocity that averages recent gradients. Zigzags across a valley cancel out, while consistent directions build up speed.
The race. Here they race on a narrow valley. Plain gradient descent zigzags. Adam adapts its step size for each direction and heads almost straight for the minimum. Momentum accelerates along the valley and ends up closest by the last step.
Adam. Adam combines momentum with a running average of squared gradients. Dividing by it gives every parameter its own step size. It works well out of the box, so Adam, or its variant AdamW, is a common default.
Recap. To recap. SGD takes cheap noisy steps. Momentum smooths the path. Adam adapts step sizes per parameter. And learning rate schedules help every optimiser.