Policy Gradients, Actor–Critic and PPO — lecture notes
Optimise the policy directly: the policy-gradient theorem and REINFORCE, baselines and advantages, actor–critic methods, PPO’s clipped objective — and how the same ideas align large language models.
0:001. Introduction

So far we learned values and derived a policy from them. In this deep dive we optimise the policy directly. Policy gradient methods handle continuous actions and stochastic strategies naturally, and one of them, proximal policy optimisation, became the workhorse of modern reinforcement learning, including the training of chat assistants.
0:212. Value-based vs policy-based

Value based methods learn Q and act greedily, which needs a maximum over a discrete set of actions and yields deterministic policies. Policy based methods learn the policy itself, a network that outputs action probabilities or the parameters of a continuous distribution. Continuous actions and stochastic strategies come for free.
0:423. A parameterised policy

The policy is a neural network with weights theta. For discrete actions, it outputs a softmax over action scores. For continuous actions, it outputs the mean and spread of a Gaussian distribution, and the action is sampled from it. Sampling gives exploration automatically.
1:004. The objective

The objective is simple: the expected return of trajectories produced by the policy. We want to climb this objective, so we use gradient ascent. The difficulty is that the return depends on the environment, which we cannot differentiate. The policy gradient theorem shows how to get around this.
1:215. The log-derivative trick

The key mathematical step is the log derivative trick: the gradient of a probability equals the probability times the gradient of its logarithm. This turns the gradient of an expected return into an expectation we can estimate from samples, and it only requires differentiating our own policy, never the environment.
1:426. The policy gradient theorem

The policy gradient theorem gives a beautiful answer. The gradient is the expected value of the return times the gradient of the log probability of the action taken. In words: push up the probability of actions in proportion to how good their outcomes were. No model of the environment is required.
2:047. Pause and think

Pause and think. In plain REINFORCE, an action is followed by a total return of zero. How does the update change that action’s probability? Not at all, because the update is the return times the log probability gradient. That reveals a problem: if all returns are positive, every action gets pushed up.
2:268. Baselines

The fix is a baseline. Subtract a reference value, such as the average return so far, from each return before using it. Actions that did better than usual get pushed up, and actions that did worse get pushed down. Remarkably, this leaves the expected gradient unchanged while cutting its variance dramatically.
2:479. REINFORCE with a baseline

Here is REINFORCE with a baseline on a problem with four actions whose average rewards are point two, point five, point nine and point four. Each update samples an action, compares its reward with the running baseline, and shifts probability accordingly. After a hundred and sixty updates, action three has about ninety seven percent of the probability.
3:1110. Pause and think

Pause and think. With a softmax policy over three actions, an update makes action two more likely. What must happen to actions one and three? Their probabilities must fall in total, because the three always sum to one. Pushing one action up automatically pushes the others down.
3:3111. Continuous actions

For continuous actions, such as a joint torque, the policy outputs a Gaussian. Early in training it is wide, exploring many torques. As learning finds that a torque near one point six works best, the mean shifts towards it and the spread narrows, just as this curve moves and sharpens.
3:5212. Advantage

The best baseline is the state’s value, and return minus value is called the advantage: how much better an action turned out than the policy’s average behaviour in that state. A positive advantage makes the action more likely, a negative one less likely. Advantages are central to all modern policy methods.
4:1413. Actor–critic

Actor critic methods learn two networks. The actor is the policy that chooses actions. The critic is a value function that judges states. After each step, the critic computes the TD error, which serves as an estimate of the advantage. The actor is nudged by it, and the critic is updated to predict better.
4:3714. Why actor–critic?

This is the same bias variance trade off we met with TD learning. REINFORCE uses full returns, which are unbiased but noisy. Actor critic bootstraps from the critic, trading a little bias for much lower variance and online learning. Generalised advantage estimation blends the two with a parameter lambda.
4:5815. Step size danger

Policy methods have a dangerous failure mode. In supervised learning the data is fixed, so a bad step can be undone. In reinforcement learning the policy generates its own data, so one overly large update can wreck performance, after which the agent collects poor experience and struggles to recover.
5:1816. Trust regions

Trust region policy optimisation, TRPO, addressed this by limiting how much each update can change the policy, measured by the KL divergence between the old and new policies. It worked well but required complex second order optimisation, which motivated a simpler alternative.
5:3617. The probability ratio

Proximal policy optimisation, PPO, starts from a probability ratio: how much more likely the new policy makes an action compared with the old policy that collected the data. A ratio of one means no change, and one point five means the action became fifty percent more likely.
5:5718. The clipped objective

PPO multiplies the ratio by the advantage, but clips the ratio to stay between point eight and one point two. For a good action, green, the objective stops rising once the action is twenty percent more likely. For a bad action, red, it stops rewarding further decreases below point eight. Big, risky policy jumps earn nothing extra.
6:2019. Pause and think

Pause and think. An action has a positive advantage, and its probability ratio has already reached one point three five. What gradient does the clipped objective give for this sample? Zero. The ratio is beyond one point two, where the objective is flat, so this sample stops pushing the policy further.
6:4220. The PPO loop

In practice PPO runs a simple loop. Collect a few thousand steps with the current policy, often in many parallel environments. Compute advantages with the critic. Run several epochs of minibatch gradient updates on the clipped objective, plus a value loss and an entropy bonus. Then discard the data and repeat.
7:0421. PPO hyperparameters

The original PPO paper used remarkably stable settings for continuous control: a clip range of point two, two thousand and forty eight steps of experience per update, ten epochs of minibatch updates with minibatches of sixty four, a learning rate of three times ten to the minus four, gamma of point nine nine and a GAE lambda of point nine five.
7:2822. Pause and think

Pause and think. PPO reuses each batch for about ten epochs and then throws it away. Why not keep it in a replay buffer, like DQN? Because PPO is on policy: its gradient assumes the data came from roughly the current policy. Old data comes from a different policy, and the ratio correction breaks down.
7:5123. A trained policy

Here is what a well trained policy looks like on CartPole. It reads the cart position and velocity and the pole angle and angular velocity, and pushes left or right to keep the pole balanced, earning plus one every step. A random policy fell after about twenty three steps. This one keeps going indefinitely.
8:1424. Why PPO won

PPO became the default because it is simple, needing only ordinary gradient descent, and robust across a wide range of tasks with both discrete and continuous actions. Its weaknesses are that it is on policy, throwing data away after a few epochs, so it is sample hungry, and many small implementation details matter.
8:3625. PPO and language models

The same algorithm aligns language models. In reinforcement learning from human feedback, people compare model answers, a reward model learns their preferences, and PPO fine tunes the language model to produce answers the reward model scores highly. Each generated token is an action, and the full reply earns the reward.
8:5726. The KL leash

There is one extra ingredient: a KL penalty that keeps the tuned model close to the original. Without this leash, the policy can exploit weaknesses in the reward model, producing strange text that scores well but reads badly. The penalty strength beta balances preference against staying natural.
9:1727. Search plus learning

Policy and value networks also power AlphaZero. During play it runs a tree search in which the policy network suggests promising moves and the value network judges positions, instead of searching every branch. Trained purely by self play, it mastered chess, shogi and Go from the rules alone.
9:3828. Beyond PPO

PPO is not the end of the story. Soft actor critic and TD3 are off policy actor critic methods popular in robotics. AlphaZero combines search with policy and value networks. Direct preference optimisation aligns models without a reinforcement learning loop, and group relative methods train reasoning models. The core ideas reappear everywhere.
10:0029. PPO loss in code

The PPO loss is only a few lines. Compute the probability ratio from log probabilities. Take the minimum of the unclipped and clipped terms, a pessimistic choice, and negate it because optimisers minimise. Then add the critic’s value loss and subtract a small entropy bonus to keep exploring.
10:2030. Pause and think

Pause and think. Why does the PPO loss include an entropy bonus? Entropy measures how random the policy is. Rewarding it slightly keeps the agent exploring, and prevents the policy from collapsing too early onto a single action before it has found the best one.
10:3931. Recap

To recap. Policy gradient methods push up the log probability of actions in proportion to their returns. Baselines and advantages reduce variance. Actor critic methods use the critic’s TD error to train the actor. PPO clips the probability ratio to keep updates safe, and RLHF uses PPO with a KL leash to align language models.
Key takeaways
- Policy-based methods learn π(a | s; θ) directly and handle continuous actions and stochastic policies naturally.
- REINFORCE: ∇J = E[Gₜ ∇log π(aₜ | sₜ)] — make actions with high returns more likely.
- Subtracting a baseline (ideally V(s)) keeps the gradient unbiased but reduces variance; G − V is the advantage.
- Actor–critic methods use a learned critic’s TD error as the advantage signal.
- PPO clips the ratio π_new/π_old to [1 − ε, 1 + ε] (ε = 0.2) to prevent destructive updates.
- RLHF fine-tunes language models with PPO against a reward model plus a KL penalty to the reference model.
Check yourself
- In REINFORCE, what multiplies ∇log π(a | s)?
Show answer
The return (or advantage) that followed the action — Good outcomes push their actions up.
- Why subtract a baseline from the return?
Show answer
To reduce the variance of the gradient without biasing it — The expected gradient is unchanged; variance falls.
- In actor–critic, what does the critic learn?
Show answer
A value function used to judge the actor’s actions — Its TD error estimates the advantage.
- PPO’s clipping keeps the probability ratio within…
Show answer
[1 − ε, 1 + ε] — Typically ε = 0.2, so [0.8, 1.2].
- In RLHF, what is the KL penalty for?
Show answer
Keeping the tuned model close to the reference model to prevent reward hacking — It is a leash on how far the policy drifts.
Go deeper
- Policy Gradient Methods: REINFORCE and the Policy Gradient Theorem · The AI Lecture Hall
- Actor–Critic Methods: A2C, A3C and Advantage Estimation · The AI Lecture Hall
- PPO and TRPO: Stable Policy Optimisation with Trust Regions · The AI Lecture Hall
- RLHF: Reinforcement Learning from Human Feedback · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/policy-gradients-and-ppo.html