AI in Motion

Backpropagation, Step by Step

Deep LearningIntermediate1:406 chapters

The chain rule in action: compute values forward, then pass gradients backward through a tiny computational graph — with real numbers.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 In the example, what is ∂f/∂z?
Q2 What does an addition node do to the incoming gradient?
Q3 Backpropagation is mainly an application of…

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. A network may have millions of weights. After a mistake, how do we know how to change each one? Backpropagation answers that efficiently, using the chain rule from calculus.

A tiny example. Take f equals x plus y, times z, with x equals minus 2, y equals 5 and z equals minus 4. Forward: q equals x plus y equals 3, and f equals q times z equals minus 12. Now backward. Start with a gradient of 1 at f. At the multiply node, the gradient for z is q, which is 3, and for q it is z, minus 4. The add node passes minus 4 to both x and y.

Chain rule. That is the chain rule. The gradient of f with respect to x equals the gradient of f with respect to q, times the gradient of q with respect to x. Minus four times one gives minus four. Backpropagation applies this rule node by node, from the output back to the inputs.

In a real network. In a real network it is the same idea at scale. The forward pass computes predictions and the loss. The backward pass sends error signals back through every layer, giving each weight a gradient. Then gradient descent updates them all.

Why it matters. Backpropagation is efficient. One backward pass gives every weight its gradient, at roughly the cost of a forward pass. Modern frameworks do it automatically, a feature called automatic differentiation.

Recap. To recap. Compute values forward. Apply the chain rule backward. Every weight gets a gradient. And gradient descent uses those gradients to improve the network.