AI in Motion

Deep LearningIntermediate1:40 video6 chapters

Backpropagation, Step by Step — lecture notes

The chain rule in action: compute values forward, then pass gradients backward through a tiny computational graph — with real numbers.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Backpropagation, Step by Step

A network may have millions of weights. After a mistake, how do we know how to change each one? Backpropagation answers that efficiently, using the chain rule from calculus.

0:122. A tiny example

A tiny example — Backpropagation, Step by Step

Take f equals x plus y, times z, with x equals minus 2, y equals 5 and z equals minus 4. Forward: q equals x plus y equals 3, and f equals q times z equals minus 12. Now backward. Start with a gradient of 1 at f. At the multiply node, the gradient for z is q, which is 3, and for q it is z, minus 4. The add node passes minus 4 to both x and y.

0:363. Chain rule

Chain rule — Backpropagation, Step by Step

That is the chain rule. The gradient of f with respect to x equals the gradient of f with respect to q, times the gradient of q with respect to x. Minus four times one gives minus four. Backpropagation applies this rule node by node, from the output back to the inputs.

0:584. In a real network

In a real network — Backpropagation, Step by Step

In a real network it is the same idea at scale. The forward pass computes predictions and the loss. The backward pass sends error signals back through every layer, giving each weight a gradient. Then gradient descent updates them all.

1:165. Why it matters

Why it matters — Backpropagation, Step by Step

Backpropagation is efficient. One backward pass gives every weight its gradient, at roughly the cost of a forward pass. Modern frameworks do it automatically, a feature called automatic differentiation.

1:296. Recap

Recap — Backpropagation, Step by Step

To recap. Compute values forward. Apply the chain rule backward. Every weight gets a gradient. And gradient descent uses those gradients to improve the network.

Key takeaways

  • Backpropagation applies the chain rule from the output back to every weight.
  • In f = (x + y)·z with x = −2, y = 5, z = −4: f = −12, ∂f/∂z = 3, ∂f/∂x = ∂f/∂y = −4.
  • One backward pass costs about as much as a forward pass.
  • Frameworks compute gradients automatically (autodiff).

Check yourself

  1. In the example, what is ∂f/∂z?
    Show answer

    3 — f = q·z, so ∂f/∂z = q = 3.

  2. What does an addition node do to the incoming gradient?
    Show answer

    Passes it unchanged to each input — The local gradient of x + y with respect to each input is 1.

  3. Backpropagation is mainly an application of…
    Show answer

    The chain rule — It chains local derivatives together.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/backpropagation.html