AI in Motion

Reinforcement LearningDeep diveIntermediate10:55 video30 chapters

Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning — lecture notes

Model-free reinforcement learning: estimate values from sampled episodes, bootstrap with temporal-difference updates, and compare on-policy SARSA with off-policy Q-learning on the famous cliff.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Dynamic programming needed a perfect map of the world. In this deep dive the map is gone. The agent must learn purely from experience, by acting and observing what happens. These model free methods, Monte Carlo, temporal difference learning, SARSA and Q learning, are the heart of practical reinforcement learning.

0:212. Model-free learning

Model-free learning — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Model free reinforcement learning learns values or policies directly from samples of experience, without ever knowing the transition probabilities. A child learns to ride a bicycle without writing down the physics. The price is that we need many samples, and we must explore to get them.

0:403. Monte Carlo prediction

Monte Carlo prediction — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

The simplest idea is Monte Carlo prediction. Play complete episodes, and for each state you visited, write down the return that actually followed. Average these returns over many episodes and you have an estimate of the state’s value, just as averaging many restaurant visits estimates its quality.

1:004. Incremental averaging

Incremental averaging — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

In practice we keep a running average. After each episode, nudge the estimate a step of size alpha towards the observed return. The term in brackets is the error: what actually happened minus what we expected. This pattern, estimate plus step size times error, appears in every method in this lecture.

1:225. The waiting problem

The waiting problem — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Monte Carlo has two weaknesses. It must wait until the end of the episode to know the return, so a very long or never ending task gives no learning at all along the way. And returns are noisy, because they depend on every random event that happens afterwards.

1:426. Pause and think

Pause and think — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Pause and think. A thermostat controller runs forever, so there are no episodes at all. Can basic Monte Carlo learn its values? No, because it waits for an episode to end before it can compute a return. We need a method that learns from each step as it happens.

2:037. Temporal-difference learning

Temporal-difference learning — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Temporal difference learning fixes both. Instead of waiting for the final return, it updates after every single step, using the reward just received plus its current guess of the next state’s value. Learning a guess from a guess is called bootstrapping. It is like updating your arrival time estimate at every traffic light.

2:268. The TD(0) update

The TD(0) update — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

The TD update replaces the full return with a one step target: the reward plus the discounted value of the next state. The difference between this target and the current estimate is called the TD error, delta. It measures the surprise of a single step, and it may be how dopamine neurons signal reward prediction errors in the brain.

2:509. A TD update by hand

A TD update by hand — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Let us do a TD update by hand. The current state is worth point five and the next state point seven. The reward is zero, gamma is one and alpha is point one. The TD error is point seven minus point five, which is point two, so the new value is point five plus point zero two, which is point five two.

3:1410. TD versus Monte Carlo

TD versus Monte Carlo — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Here is the classic comparison on a five state random walk. Episodes start in the middle state, C, and step left or right at random. Exiting on the right gives one, and the true values are one sixth to five sixths. Averaged over a hundred runs, TD’s error drops faster and lower than Monte Carlo’s.

3:3711. Bias and variance

Bias and variance — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Why does TD often win? The Monte Carlo target is the real return, so it is unbiased, but it has high variance. The TD target depends on only one random step, so its variance is much lower, although it is slightly biased by the current guess. In many problems the lower variance matters more.

4:0012. Pause and think

Pause and think — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Pause and think. After a single step, the TD error is large and positive. What happened? The outcome was better than expected: the reward plus the next state’s value exceeded our estimate. So the update raises the value of the state we just left, moving it towards that better target.

4:2113. From values to control

From values to control — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

So far we have predicted values for a fixed behaviour. To improve behaviour without a model, we learn action values, Q of s and a, instead of state values. Knowing V alone would not tell us which action to take unless we knew where each action leads. With Q, we simply pick the best action.

4:4414. ε-greedy

ε-greedy — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

To keep exploring, agents usually act epsilon greedily. Most of the time they take the action with the highest Q value, but with a small probability epsilon, say one step in ten, they try a random action. Without this, an agent could lock onto a mediocre habit and never discover anything better.

5:0615. SARSA

SARSA — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

SARSA is TD learning for action values. Its target uses the action the agent will actually take next, which might be an exploratory one. Its name comes from the five quantities it uses: state, action, reward, next state and next action. SARSA learns the value of the policy it is really following, exploration included.

5:2916. Q-learning

Q-learning — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Q learning, introduced by Chris Watkins in 1989, makes one change. Its target assumes the best possible next action, the maximum over Q, regardless of what the agent actually does next. So it learns the optimal action values, even while it behaves exploratorily. This is called off policy learning.

5:5017. Q-learning in the grid

Q-learning in the grid — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Here is Q learning in our grid world, with deterministic moves this time. Each cell shows four triangles, one Q value per action, with green for good and red for bad. The first episode stumbles around for a hundred and twenty seven steps. By episode forty, the greedy path to the goal takes only about twelve.

6:1318. On-policy vs off-policy

On-policy vs off-policy — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

This distinction matters. An on policy method, like SARSA, learns about the policy it is following. An off policy method, like Q learning, learns about a different target policy, usually the greedy one, from exploratory behaviour. Off policy learning can even reuse old experience or someone else’s demonstrations.

6:3319. Cliff walking

Cliff walking — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

The famous cliff walking example shows the difference. Stepping into the cliff costs a hundred and sends you back to the start. Q learning finds the optimal path right along the edge, but its random exploratory steps keep knocking it off during training. SARSA accounts for its own exploration and learns the safer path along the top.

6:5720. Pause and think

Pause and think — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Pause and think. Q learning finds the shorter path, yet SARSA earns more reward per episode while training, about minus twenty eight versus minus forty five. Why? Q learning’s path hugs the edge, and with ten percent random exploration it keeps falling off. SARSA values the path it really follows, so it stays safe.

7:2021. Which is better?

Which is better? — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Which is better? It depends. If exploration stops after training, Q learning’s greedy policy is optimal. If a real robot must keep exploring while operating, SARSA’s caution may be wiser. And if epsilon is slowly decayed to zero, both methods converge to the optimal policy.

7:3922. Expected SARSA

Expected SARSA — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

A middle ground is expected SARSA. Instead of using the single next action that was sampled, or the maximum, it uses the average next value under the current policy. This removes the randomness of the next action choice, lowering variance, and with a greedy target policy it becomes exactly Q learning.

8:0123. Monte Carlo control

Monte Carlo control — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Monte Carlo methods can also control. Average returns to estimate Q, act greedily with some exploration, and repeat, which is generalised policy iteration again. To make sure every action keeps being tried, they use epsilon soft policies or start episodes from random state action pairs, called exploring starts.

8:2124. n-step methods

n-step methods — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Monte Carlo and TD are two ends of a spectrum. An n step method uses n real rewards before bootstrapping from an estimate. One step is TD zero, and infinitely many steps is Monte Carlo. TD lambda blends all of them with decaying weights. Intermediate choices often learn fastest in practice.

8:4325. Q-learning in code

Q-learning in code — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

The code is short. Keep a table of Q values starting at zero. In every step, choose an action epsilon greedily, take it, and observe the reward and next state. The target is the reward plus the discounted best next Q value, and the table entry moves a step alpha towards it.

9:0526. Hyperparameters

Hyperparameters — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

There are only a few knobs. The learning rate alpha trades speed for noise. The discount gamma sets the horizon. Epsilon controls exploration and is often decayed over time. And the initial Q values matter: setting them optimistically high makes untried actions look attractive, a simple way to encourage exploration.

9:2627. Convergence

Convergence — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Tabular Q learning comes with a guarantee: it converges to the optimal action values, provided every state action pair keeps being visited and the learning rate decays appropriately. That guarantee disappears once we replace the table with a neural network, a story for the deep Q network lecture.

9:4728. Maximisation bias

Maximisation bias — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Q learning has one subtle flaw: taking the maximum of noisy estimates is biased upwards, because the maximum tends to pick values that happen to be overestimated. Double Q learning fixes this by using one table to choose the best action and a second table to evaluate it. The same trick improves deep Q networks.

10:1029. Pause and think

Pause and think — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

Pause and think. Your agent now faces a maze with a billion possible states. Will a Q table still work? No. The table would be enormous, and most states would never be visited even once. We need function approximation, such as a neural network, to generalise from states it has seen to similar ones it has not.

10:3430. Recap

Recap — Learning from Experience: Monte Carlo, TD, SARSA and Q-Learning

To recap. Monte Carlo averages complete returns, while TD learning bootstraps after every step from a one step target. SARSA is on policy and Q learning is off policy, using the maximum over next actions. The cliff shows their different behaviour, and exploration, step size and discount are the key knobs.

Key takeaways

  • Model-free methods learn from sampled experience without knowing transition probabilities.
  • Monte Carlo uses full returns (unbiased, high variance); TD uses r + γV(s′) (bootstrapped, lower variance).
  • The TD error δ = r + γV(s′) − V(s) drives learning after every step.
  • SARSA (on-policy) uses the next action actually taken; Q-learning (off-policy) uses maxₐ′ Q(s′, a′).
  • In cliff walking, Q-learning learns the optimal edge path but SARSA earns more reward during ε-greedy training.
  • Tables do not scale; large problems need function approximation.

Check yourself

  1. What is the TD(0) target for V(s)?
    Show answer

    r + γV(s′) — One real reward plus a bootstrapped estimate.

  2. Which method must wait until the end of an episode to update?
    Show answer

    Monte Carlo — Monte Carlo needs the complete return.

  3. What makes Q-learning off-policy?
    Show answer

    Its target uses the max over next actions, not the action actually taken — It learns the greedy policy while behaving exploratorily.

  4. On the cliff with ε = 0.1, which method earns more reward during training?
    Show answer

    SARSA — SARSA learns the safer path, accounting for its own exploration.

  5. Double Q-learning addresses…
    Show answer

    Overestimation from taking the max of noisy estimates — It decouples action selection and evaluation.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/monte-carlo-and-td-learning.html