Backpropagation, Step by Step — lecture notes
The chain rule in action: compute values forward, then pass gradients backward through a tiny computational graph — with real numbers.
0:001. Introduction

A network may have millions of weights. After a mistake, how do we know how to change each one? Backpropagation answers that efficiently, using the chain rule from calculus.
0:122. A tiny example

Take f equals x plus y, times z, with x equals minus 2, y equals 5 and z equals minus 4. Forward: q equals x plus y equals 3, and f equals q times z equals minus 12. Now backward. Start with a gradient of 1 at f. At the multiply node, the gradient for z is q, which is 3, and for q it is z, minus 4. The add node passes minus 4 to both x and y.
0:363. Chain rule

That is the chain rule. The gradient of f with respect to x equals the gradient of f with respect to q, times the gradient of q with respect to x. Minus four times one gives minus four. Backpropagation applies this rule node by node, from the output back to the inputs.
0:584. In a real network

In a real network it is the same idea at scale. The forward pass computes predictions and the loss. The backward pass sends error signals back through every layer, giving each weight a gradient. Then gradient descent updates them all.
1:165. Why it matters

Backpropagation is efficient. One backward pass gives every weight its gradient, at roughly the cost of a forward pass. Modern frameworks do it automatically, a feature called automatic differentiation.
1:296. Recap

To recap. Compute values forward. Apply the chain rule backward. Every weight gets a gradient. And gradient descent uses those gradients to improve the network.
Key takeaways
- Backpropagation applies the chain rule from the output back to every weight.
- In f = (x + y)·z with x = −2, y = 5, z = −4: f = −12, ∂f/∂z = 3, ∂f/∂x = ∂f/∂y = −4.
- One backward pass costs about as much as a forward pass.
- Frameworks compute gradients automatically (autodiff).
Check yourself
- In the example, what is ∂f/∂z?
Show answer
3 — f = q·z, so ∂f/∂z = q = 3.
- What does an addition node do to the incoming gradient?
Show answer
Passes it unchanged to each input — The local gradient of x + y with respect to each input is 1.
- Backpropagation is mainly an application of…
Show answer
The chain rule — It chains local derivatives together.
Go deeper
- Backpropagation Derived Step by Step · The AI Lecture Hall
- Computational Graphs and Automatic Differentiation · The AI Lecture Hall
- Derivatives, Gradients and the Chain Rule · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/backpropagation.html