AI in Motion

Reinforcement LearningDeep diveAdvanced11:07 video32 chapters

Deep Q-Networks: Reinforcement Learning Meets Deep Learning — lecture notes

How DQN learned Atari from pixels: function approximation, the deadly triad, experience replay, target networks, and the improvements that became Rainbow.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

In 2013 and 2015, a team at DeepMind showed that a single algorithm could learn to play dozens of Atari games directly from screen pixels, reaching human level on many of them. That algorithm was the deep Q network. In this deep dive we see exactly what made it work.

0:212. Why tables fail

Why tables fail — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Tabular Q learning keeps one number per state and action. That is hopeless here. An Atari screen is two hundred and ten by one hundred and sixty pixels, so the number of possible screens is astronomical, and robots have continuous states. We need a function that generalises from states it has seen to new ones.

0:443. Function approximation

Function approximation — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

The answer is function approximation. Instead of a table, we use a parameterised function, Q of s and a with weights theta, such as a neural network. It is trained to match TD targets, and similar states automatically get similar values, so experience generalises.

1:034. The DQN architecture

The DQN architecture — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Here is the DQN architecture. The state is a stack of the last four frames, so motion is visible. Convolutional layers extract features, a dense layer combines them, and the output layer has one Q value per action. A single forward pass scores every action, and the agent picks the highest, here up.

1:255. Generalisation

Generalisation — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

A neural network generalises because states that look alike produce similar activations, and therefore similar Q values. Learning that a ball approaching the paddle is dangerous in one position automatically transfers to nearby positions the agent has never seen. That is the power, and the risk, of function approximation.

1:466. Preprocessing

Preprocessing — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

DQN also depended on careful preprocessing. Frames were converted to greyscale and shrunk to eighty four by eighty four pixels. Each chosen action was repeated for four frames, and the last four processed frames were stacked as the state. This kept computation manageable while preserving the information needed to play.

2:077. The loss

The loss — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Training turns Q learning’s update into a regression problem. The loss is the squared difference between the TD target, reward plus discounted best next Q value, and the network’s current prediction. Gradient descent on this loss nudges the network towards consistent Q values.

2:268. Huber loss

Huber loss — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

In practice the squared error is often replaced with the Huber loss, which is quadratic for small errors but linear for large ones. A few enormous TD errors early in training then cannot produce enormous gradients. The original DQN achieved the same effect by clipping the error term between minus one and one.

2:489. The deadly triad

The deadly triad — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Unfortunately, combining function approximation, bootstrapping and off policy learning, the so called deadly triad, can make value estimates diverge. Naive neural Q learning was notoriously unstable. The DQN paper introduced two stabilisers that made the difference: experience replay and a target network.

3:0610. Two sources of instability

Two sources of instability — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

There are two main sources of instability. First, consecutive frames are nearly identical, so training on them in order is like studying one page over and over, and the network forgets everything else. Second, the target is computed by the very network being trained, so every update moves the goalposts.

3:2711. Experience replay

Experience replay — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Experience replay fixes the first problem. Every transition, state, action, reward and next state, is stored in a large buffer, up to a million of them in the original DQN. Training samples random minibatches from the buffer, which breaks the correlations and lets each experience be reused many times.

3:4812. The target network

The target network — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

A target network fixes the second problem. We keep a frozen copy of the Q network, and use it only to compute targets. Every ten thousand steps in the original DQN, the frozen copy is refreshed with the current weights. Between refreshes the target stays still, so learning is far more stable.

4:1013. Pause and think

Pause and think — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Pause and think. Why does sampling random minibatches from a replay buffer help, compared with training on the latest transitions in order? Random samples are closer to independent and cover many past situations, so each gradient step is less biased, and the network does not forget what it learned earlier.

4:3114. The DQN algorithm

The DQN algorithm — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Here is the whole algorithm. Observe a stack of frames. Act epsilon greedily, with epsilon decaying from one to about point one. Store the transition in the replay buffer. Sample a random minibatch, compute targets with the frozen network, and take a gradient step. Periodically copy the weights into the target network.

4:5315. Exploration schedule

Exploration schedule — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Exploration in DQN followed a schedule. Epsilon started at one, meaning completely random actions, and fell linearly to point one over the first million frames, then stayed there. Early random play fills the replay buffer with diverse experience before the network’s own judgement takes over.

5:1316. Pause and think

Pause and think — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Pause and think. DQN clipped every reward to minus one, zero or plus one in all games. What does that gain, and what does it lose? One set of hyperparameters can work across games whose scores differ enormously. But the agent can no longer tell a small reward from a huge one.

5:3517. Results

Results — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

The 2015 Nature paper evaluated forty nine Atari games with a single algorithm and one set of hyperparameters, learning from pixels and the score alone. It used four stacked frames, a replay buffer of one million transitions, and about fifty million frames of training per game, reaching human level performance on many of the games.

5:5818. Pause and think

Pause and think — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Pause and think. Why did DQN stack four frames instead of using just one as the state? Because a single frame hides motion: you cannot tell which way the ball is moving. Stacking a few frames restores that information and makes the state approximately Markov.

6:1719. Double DQN

Double DQN — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Researchers soon improved DQN. Double DQN tackles the overestimation caused by taking a maximum over noisy estimates. It uses the online network to choose the best next action and the target network to evaluate it. This simple change gave more accurate values and better scores.

6:3620. Why max overestimates

Why max overestimates — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Why does the maximum overestimate? Suppose four actions are all truly worth zero, but each estimate has random noise of about plus or minus one. The maximum of four noisy estimates is about plus one on average, not zero. The max systematically picks lucky overestimates, which is the bias Double DQN removes.

6:5821. Dueling networks

Dueling networks — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

The dueling architecture splits the output into two streams: one estimates how good the state is overall, and the other estimates the advantage of each action. In many states the action barely matters, like when no enemy is on screen, and learning the state value directly makes training more efficient.

7:1922. Prioritised replay

Prioritised replay — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Prioritised experience replay samples surprising transitions, those with large TD errors, more often than boring ones, with a correction to avoid bias. It is like revising the exam questions you got wrong rather than the ones you already know. It made learning noticeably faster.

7:3823. Rainbow

Rainbow — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Rainbow combined six of these improvements in one agent: double Q learning, prioritised replay, dueling networks, multi step returns, distributional value learning and noisy networks for exploration. Together they were far more data efficient than the original DQN, and each component contributed.

7:5624. Distributional RL

Distributional RL — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

One of those components deserves a closer look. Distributional reinforcement learning predicts the full distribution of possible returns rather than just their average. A guaranteed ten and a coin flip between zero and twenty have the same mean but very different risks, and representing that difference improves learning.

8:1725. DQN in code

DQN in code — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Here is the core update in PyTorch. Sample a minibatch from the buffer. Pick out the Q values of the actions that were taken. Compute targets with the frozen target network, without gradients. Take a gradient step on a robust regression loss, and every ten thousand steps copy the weights into the target network.

8:4026. A replay buffer

A replay buffer — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

A replay buffer is simple to build. A fixed size queue stores transitions, and when it is full the oldest ones fall out. Every step pushes the newest transition, and training samples a uniformly random minibatch. Prioritised replay replaces the uniform sampling with sampling weighted by TD error.

9:0027. Offline RL

Offline RL — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Taking replay to the extreme gives offline reinforcement learning: learning a policy purely from a fixed dataset of logged experience, with no new interaction at all. This matters in healthcare or robotics, where exploration is costly or dangerous. The key challenge is not trusting Q values for actions the data never tried.

9:2228. Limitations

Limitations — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Value based deep reinforcement learning has limits. The maximum over actions needs a small, discrete set of actions, so continuous control, like robot joint torques, needs other methods such as DDPG, TD3 and SAC. DQN is also sample hungry, needing tens of millions of frames, and results can be sensitive to settings and random seeds.

9:4529. Pause and think

Pause and think — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

Pause and think. A robot arm chooses any torque between minus two and plus two for each of seven joints. Why is plain DQN awkward? It needs the maximum over a finite list of actions. Here there are infinitely many, and discretising all seven joints explodes combinatorially. Actor critic methods handle this naturally.

10:0730. Evaluating agents

Evaluating agents — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

How is progress measured? Atari results are usually reported as human normalised scores: zero percent means random play and one hundred percent means the human reference, summarised by the median across games. Because results vary so much between runs, careful papers report several random seeds.

10:2731. Legacy

Legacy — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

DQN launched modern deep reinforcement learning. Rainbow combined its improvements, distributed versions ran hundreds of actors feeding one replay buffer, MuZero learned its own model of the game and planned with it, and in 2020 Agent57 became the first agent to beat the human benchmark on all fifty seven Atari games.

10:4832. Recap

Recap — Deep Q-Networks: Reinforcement Learning Meets Deep Learning

To recap. DQN replaces the Q table with a neural network trained on TD targets. Experience replay breaks correlations and reuses data, and a frozen target network keeps targets stable. Double Q learning, dueling heads, prioritised replay and distributional learning combine into Rainbow.

Key takeaways

  • Deep Q-networks approximate Q(s, a) with a neural network so learning generalises across huge state spaces.
  • The deadly triad (function approximation + bootstrapping + off-policy) can cause divergence.
  • Experience replay samples random minibatches from a large buffer (1M transitions in DQN).
  • A target network, synced every C steps (10,000 in DQN), stabilises the TD targets.
  • Double DQN, dueling networks, prioritised replay, multi-step returns, distributional RL and noisy nets combine into Rainbow.
  • DQN needs discrete actions; continuous control uses actor–critic methods such as DDPG, TD3 and SAC.

Check yourself

  1. What problem does experience replay mainly address?
    Show answer

    Highly correlated consecutive training samples — Random minibatches break temporal correlations.

  2. What is the target network for?
    Show answer

    Providing stable TD targets that do not shift with every update — It is a frozen copy refreshed periodically.

  3. Double DQN reduces…
    Show answer

    Overestimation of Q-values — Choosing and evaluating with different networks removes max bias.

  4. Why did DQN stack 4 frames as its state?
    Show answer

    To reveal motion so the state is closer to Markov — A single frame hides velocity.

  5. Which setting is plain DQN poorly suited to?
    Show answer

    Continuous robot joint torques — The max over actions needs a finite action set.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/deep-q-networks.html