Gradient Descent and the Learning Rate — lecture notes
How almost every ML model learns: follow the slope downhill. See small, good and too-large learning rates, and how optimisers like Adam move on a loss surface.
0:001. Introduction

Almost every machine learning model, from linear regression to giant neural networks, learns with gradient descent. The idea is to walk downhill on the error.
0:112. Following the slope

The curve shows the loss for different values of one weight. We measure the slope, the dashed line, and step the opposite way. Where it is steep, steps are big. Near the bottom, the slope flattens and steps shrink automatically.
0:283. Update rule

The update rule is short. The new weight equals the old weight, minus the learning rate times the gradient. The minus sign means we always move downhill.
0:404. Too small

With a learning rate that is too small, every step is tiny. After fourteen steps, the weight has only crawled to about 2.6, far from the minimum at 5.5. Training would take forever.
0:555. Too large

With a learning rate that is too large, each step overshoots the valley. The weight bounces from one side to the other and never settles. The loss does not go down.
1:096. Optimisers

Real models have many weights, so the loss is a surface. On this narrow valley, plain gradient descent zigzags from side to side. Adam adapts its step size for each direction and heads almost straight in. Momentum builds up speed along the valley, and by the end it is closest to the minimum.
1:317. Recap

To recap. Gradient descent steps against the slope. The learning rate sets the step size, neither too small nor too large. And optimisers like momentum and Adam make the journey faster and smoother.
Key takeaways
- Gradient descent updates w ← w − η·∂L/∂w.
- A small learning rate is slow; a large one overshoots and may diverge.
- Steps naturally shrink as the slope flattens near a minimum.
- Momentum and Adam improve on plain gradient descent in narrow valleys.
Check yourself
- What happens if the learning rate is far too large?
Show answer
Steps overshoot and the loss may never settle — Large steps jump over the minimum repeatedly.
- Why do steps get smaller near the minimum?
Show answer
The gradient (slope) gets smaller — Step size is learning rate × gradient, and the slope flattens.
- Which optimiser adapts its step size separately for each direction?
Show answer
Adam — Adam scales updates using running estimates of gradient size.
Go deeper
- Gradient Descent: Theory, Step Sizes and Convergence · The AI Lecture Hall
- Optimisers I: SGD, Momentum and Nesterov Acceleration · The AI Lecture Hall
- Optimisers II: AdaGrad, RMSProp, Adam and AdamW · The AI Lecture Hall
- Learning Rate Schedules, Warm-up and the LR Range Test · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/gradient-descent.html