AI in Motion
♞

Reinforcement Learning

MDPs and returns, value and policy iteration, Monte Carlo and TD learning, Q-learning vs SARSA, bandits, deep Q-networks, policy gradients and PPO.

6 animated lectures · 66 minutes

▶ Start with lecture 1
Deep dive11:04
Reinforcement Learning · 01

Reinforcement Learning Foundations: Agents, MDPs and Returns

How an agent learns from rewards: the agent–environment loop, Markov decision processes, discounted returns, policies and value functions — the vocabulary behind every RL algorithm.

Beginner · 31 chapters · 5-question quiz
Deep dive11:12
Reinforcement Learning · 02

Dynamic Programming: Value Iteration and Policy Iteration

When the rules of the world are known, the Bellman equations can be solved exactly. Watch value iteration spread value from the goal and policy iteration converge in five rounds.

Intermediate · 32 chapters · 5-question quiz
Deep dive10:55
Reinforcement Learning · 03

Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Model-free reinforcement learning: estimate values from sampled episodes, bootstrap with temporal-difference updates, and compare on-policy SARSA with off-policy Q-learning on the famous cliff.

Intermediate · 30 chapters · 5-question quiz
Deep dive10:59
Reinforcement Learning · 04

Exploration and Multi-Armed Bandits

The purest form of the explore–exploit dilemma: greedy, ε-greedy, UCB and Thompson sampling, regret, and where bandits run in the real world — from A/B tests to recommendations.

Intermediate · 31 chapters · 5-question quiz
Deep dive11:07
Reinforcement Learning · 05

Deep Q-Networks: Reinforcement Learning Meets Deep Learning

How DQN learned Atari from pixels: function approximation, the deadly triad, experience replay, target networks, and the improvements that became Rainbow.

Advanced · 32 chapters · 5-question quiz
Deep dive11:02
Reinforcement Learning · 06

Policy Gradients, Actor–Critic and PPO

Optimise the policy directly: the policy-gradient theorem and REINFORCE, baselines and advantages, actor–critic methods, PPO’s clipped objective — and how the same ideas align large language models.

Advanced · 31 chapters · 5-question quiz