Reinforcement Learning
MDPs and returns, value and policy iteration, Monte Carlo and TD learning, Q-learning vs SARSA, bandits, deep Q-networks, policy gradients and PPO.
6 animated lectures · 66 minutes
▶ Start with lecture 1Reinforcement Learning Foundations: Agents, MDPs and Returns
How an agent learns from rewards: the agent–environment loop, Markov decision processes, discounted returns, policies and value functions — the vocabulary behind every RL algorithm.
Dynamic Programming: Value Iteration and Policy Iteration
When the rules of the world are known, the Bellman equations can be solved exactly. Watch value iteration spread value from the goal and policy iteration converge in five rounds.
Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning
Model-free reinforcement learning: estimate values from sampled episodes, bootstrap with temporal-difference updates, and compare on-policy SARSA with off-policy Q-learning on the famous cliff.
Exploration and Multi-Armed Bandits
The purest form of the explore–exploit dilemma: greedy, ε-greedy, UCB and Thompson sampling, regret, and where bandits run in the real world — from A/B tests to recommendations.
Deep Q-Networks: Reinforcement Learning Meets Deep Learning
How DQN learned Atari from pixels: function approximation, the deadly triad, experience replay, target networks, and the improvements that became Rainbow.
Policy Gradients, Actor–Critic and PPO
Optimise the policy directly: the policy-gradient theorem and REINFORCE, baselines and advantages, actor–critic methods, PPO’s clipped objective — and how the same ideas align large language models.