AI in Motion

Reinforcement LearningDeep diveBeginner11:04 video31 chapters

Reinforcement Learning Foundations: Agents, MDPs and Returns — lecture notes

How an agent learns from rewards: the agent–environment loop, Markov decision processes, discounted returns, policies and value functions — the vocabulary behind every RL algorithm.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Reinforcement Learning Foundations: Agents, MDPs and Returns

Welcome to the first deep dive on reinforcement learning, the branch of AI where an agent learns by trial and error, guided only by rewards. It taught computers to play Atari from pixels, beat world champions at Go, and it shapes how chatbots are aligned with human preferences.

0:202. A different kind of learning

A different kind of learning — Reinforcement Learning Foundations: Agents, MDPs and Returns

Supervised learning is like studying with an answer key: every example comes with the correct label. Reinforcement learning is different. Nobody tells the agent the right action. It only receives rewards, those rewards may arrive long after the decisions that earned them, and every action changes what the agent sees next.

0:423. The agent–environment loop

The agent–environment loop — Reinforcement Learning Foundations: Agents, MDPs and Returns

Everything in reinforcement learning happens in a loop. The agent observes the state of the environment and chooses an action. The environment responds with a new state and a reward. Here each step costs minus one and reaching the star gives plus ten, so this trajectory earns a total of four.

1:034. The vocabulary

The vocabulary — Reinforcement Learning Foundations: Agents, MDPs and Returns

Let us fix the vocabulary. The state is what the agent observes, like a chess board. An action is a choice it can make, like a legal move. A reward is a number saying how good a step was. The policy is the agent’s strategy, and an episode is one complete run, like one game.

1:265. Milestones

Milestones — Reinforcement Learning Foundations: Agents, MDPs and Returns

The field has a remarkable history. In 1992 TD Gammon reached expert level at backgammon. In 2013 deep Q networks learned Atari games straight from pixels. AlphaGo beat Lee Sedol in 2016, AlphaZero mastered three board games from self play, and from 2022 reinforcement learning from human feedback shaped chat assistants.

1:486. A grid world

A grid world — Reinforcement Learning Foundations: Agents, MDPs and Returns

Here is the running example for this track: a small grid world. The agent starts in the bottom left corner. Reaching the green square gives plus one and ends the episode, falling into the red pit gives minus one, and every other step costs a little, minus point zero four, so dawdling is discouraged.

2:117. Stochastic environments

Stochastic environments — Reinforcement Learning Foundations: Agents, MDPs and Returns

This world is also slippery. When the agent tries to move up, it succeeds eighty percent of the time, but ten percent of the time it slides to the left and ten percent to the right. Real environments are like this: robots slip, markets fluctuate and opponents surprise you.

2:318. Pause and think

Pause and think — Reinforcement Learning Foundations: Agents, MDPs and Returns

Pause and think. The random agent has wandered for one hundred and sixty steps without reaching either exit. What is its total reward so far? Each step costs point zero four, so it has collected minus six point four. The small step penalty is what makes efficient paths worth learning.

2:539. The Markov property

The Markov property — Reinforcement Learning Foundations: Agents, MDPs and Returns

Reinforcement learning usually assumes the Markov property: the future depends only on the current state and action, not on the history. A chess position has this property. A single frame of a video game does not, because you cannot tell which way the ball is moving, which is why game agents stack several frames.

3:1510. Markov decision process

Markov decision process — Reinforcement Learning Foundations: Agents, MDPs and Returns

Formally, the environment is modelled as a Markov decision process, or MDP. It consists of a set of states, a set of actions, transition probabilities saying where each action might lead, a reward function, and a discount factor gamma. Almost every reinforcement learning algorithm is built on this model.

3:3611. An MDP as a diagram

An MDP as a diagram — Reinforcement Learning Foundations: Agents, MDPs and Returns

Here is a tiny MDP about a machine. Running it fast earns two points, running it slowly earns one. But running fast warms the machine half the time, and running fast while it is warm overheats it, which costs ten points and ends everything. The best policy must weigh reward now against risk later.

3:5912. Return

Return — Reinforcement Learning Foundations: Agents, MDPs and Returns

The agent’s goal is not to maximise the next reward, but the return: the sum of all future rewards from now on. A chess player happily sacrifices a piece now to win the game later. Maximising the return is what makes reinforcement learning about long term planning.

4:1913. Discounting

Discounting — Reinforcement Learning Foundations: Agents, MDPs and Returns

Usually future rewards are discounted. A reward t steps in the future is multiplied by gamma to the power t. With gamma of point nine, the plus ten reward four steps away is worth about six point five six today, and the whole return is about three point one two.

4:4014. Choosing gamma

Choosing gamma — Reinforcement Learning Foundations: Agents, MDPs and Returns

The discount factor sets how far ahead the agent looks. With gamma of one half, rewards more than a few steps away barely count. With point nine the effective horizon is about ten steps, and with point nine nine about a hundred. Discounting also keeps infinite sums finite.

5:0115. Pause and think

Pause and think — Reinforcement Learning Foundations: Agents, MDPs and Returns

Pause and think. Why would we choose gamma of point nine nine for chess rather than one half? Because chess is won or lost many moves later. With one half, the final reward would be practically invisible from the opening. A gamma close to one lets the agent value long term consequences.

5:2316. Policy

Policy — Reinforcement Learning Foundations: Agents, MDPs and Returns

The policy is the agent’s behaviour. A deterministic policy picks one action per state. A stochastic policy gives a probability for each action, which is useful for exploring and for games where being predictable is a weakness. It can be a lookup table or a large neural network.

5:4317. State value

State value — Reinforcement Learning Foundations: Agents, MDPs and Returns

How good is it to be in a state? The state value function answers this: the expected return if you start in that state and follow the policy afterwards. In our grid world, states near the goal have high value, and states near the pit have low value.

6:0418. Action value

Action value — Reinforcement Learning Foundations: Agents, MDPs and Returns

The action value function, called Q, is even more useful. It is the expected return if you take a particular action first and then follow the policy. If you know Q, acting well is easy: in every state simply pick the action with the highest Q value. That is the idea behind Q learning.

6:2619. Values in the grid world

Values in the grid world — Reinforcement Learning Foundations: Agents, MDPs and Returns

Here are state values for our grid world, computed exactly, as we will learn in the next lecture. The cell next to the goal is worth about point nine three. Values fall as we move away, reaching almost zero at the start. Brighter green means a better place to be.

6:4820. The Bellman equation

The Bellman equation — Reinforcement Learning Foundations: Agents, MDPs and Returns

Values obey a beautiful recursive relationship called the Bellman equation. The value of a state equals the expected immediate reward, plus the discounted value of the next state. Today’s value is reward now plus the discounted value of tomorrow. Nearly every algorithm in this track exploits it.

7:0821. Pause and think

Pause and think — Reinforcement Learning Foundations: Agents, MDPs and Returns

Try the Bellman equation yourself. A state gives a reward of one, and always leads to a next state worth five. With gamma equal to point nine, what is the value of the first state? It is one plus point nine times five, which equals five and a half.

7:2822. Optimality

Optimality — Reinforcement Learning Foundations: Agents, MDPs and Returns

The goal of reinforcement learning is an optimal policy: one that does at least as well as any other policy from every state. Every finite MDP has at least one. Its value functions are written V star and Q star, and finding them is the central problem of the field.

7:5023. The families of algorithms

The families of algorithms — Reinforcement Learning Foundations: Agents, MDPs and Returns

Algorithms fall into families. Model based methods know or learn how the world works and then plan. Value based methods learn Q values from experience and act greedily. Policy based methods adjust the policy directly. Actor critic methods combine a policy and a value function, and they dominate modern practice.

8:1124. Explore or exploit?

Explore or exploit? — Reinforcement Learning Foundations: Agents, MDPs and Returns

Every learning agent faces a dilemma. Should it exploit what it already knows, collecting reward now, or explore something new that might be better? Exploit too much and you never discover the best option. Explore too much and you waste time. We devote a whole lecture to this trade off.

8:3225. Credit assignment

Credit assignment — Reinforcement Learning Foundations: Agents, MDPs and Returns

Another fundamental challenge is credit assignment. When a reward finally arrives, which of the many earlier decisions deserve the credit or the blame? A chess game lost on move sixty may have been lost by a blunder on move twelve. Value functions and discounting are our main tools for sorting this out.

8:5426. Where RL is used

Where RL is used — Reinforcement Learning Foundations: Agents, MDPs and Returns

Reinforcement learning is used well beyond games. Robots learn to walk and grasp, often in simulation first. Recommender systems optimise long term engagement. It has been applied to data centre cooling, chip design and traffic control. And it is a key ingredient in aligning large language models.

9:1427. The loop in code

The loop in code — Reinforcement Learning Foundations: Agents, MDPs and Returns

In code, the loop looks the same in almost every library. Gymnasium is the standard toolkit. Create an environment, reset it to get the first state, then repeatedly choose an action and step the environment, receiving the next state and reward, until the episode ends. Here the policy is random.

9:3528. Random versus trained

Random versus trained — Reinforcement Learning Foundations: Agents, MDPs and Returns

This is CartPole, a classic benchmark: push the cart left or right to keep the pole upright, earning plus one for every step it stays up. A random policy, like the one in our code, lets the pole fall after about twenty three steps on average. By the end of this track you will know how to do far better.

9:5929. Common pitfalls

Common pitfalls — Reinforcement Learning Foundations: Agents, MDPs and Returns

Framing a problem well is half the battle. Agents are famous for reward hacking, finding loopholes that score points without doing what you meant. Sparse rewards give almost no signal. States that hide crucial information break the Markov assumption. And reinforcement learning is hungry, often needing millions of interactions.

10:2030. Pause and think

Pause and think — Reinforcement Learning Foundations: Agents, MDPs and Returns

A final question. You reward a cleaning robot plus one for every piece of dirt it picks up. What could go wrong? It might learn to tip out the dirt and pick it up again, forever. Reward the outcome you actually want, a clean room, rather than a proxy that can be gamed.

10:4231. Recap

Recap — Reinforcement Learning Foundations: Agents, MDPs and Returns

To recap. The agent and environment interact in a loop of states, actions and rewards, formalised as a Markov decision process. The goal is the discounted return, not the next reward. Value functions measure how good states and actions are, and the Bellman equation links the value now to the value next.

Key takeaways

  • RL learns from rewards through trial and error, and actions affect future states.
  • An MDP is defined by states, actions, transition probabilities, rewards and a discount factor γ.
  • The return Gₜ is the (discounted) sum of future rewards; γ sets the effective horizon ≈ 1/(1 − γ).
  • A policy maps states to actions; V(s) and Q(s, a) are expected returns under that policy.
  • The Bellman equation: V(s) = E[r + γV(s′)].
  • Main challenges: exploration vs exploitation, credit assignment, reward hacking and sample efficiency.

Check yourself

  1. With γ = 0.9, what weight does a reward 2 steps ahead get?
    Show answer

    0.81 — γ² = 0.81.

  2. What does Q(s, a) measure?
    Show answer

    The expected return after taking action a in state s and following the policy — Q is an expected return conditioned on the first action.

  3. What is the Markov property?
    Show answer

    The future depends only on the current state and action — History adds nothing beyond the present state.

  4. A state gives reward 2 and always leads to a state worth 10. With γ = 0.5, its value is…
    Show answer

    7 — 2 + 0.5 × 10 = 7.

  5. An agent finds a loophole that maximises reward without doing the intended task. This is called…
    Show answer

    Reward hacking — A classic RL failure mode.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/rl-foundations-mdps-and-returns.html