AI in Motion

Recurrent Networks and LSTMs

Deep LearningIntermediate1:326 chapters

Networks with memory: an RNN reads a sentence word by word, carrying a hidden state; an LSTM adds gates to remember for longer.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 What carries information from one time step to the next in an RNN?
Q2 Which LSTM gate decides what to erase from memory?
Q3 Why did Transformers replace RNNs for most language tasks?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. Language, music and sensor readings are sequences, where order matters. Recurrent neural networks process sequences one step at a time, carrying a memory forward.

An RNN unrolled. Here an RNN reads a sentence word by word. At each step it combines the current word with its hidden state, the memory from previous words, and produces a new hidden state. The amber arrow carries that memory forward through time.

The maths. The new hidden state combines the previous hidden state and the current input, using the same weights at every step. But long sentences cause trouble: gradients passed back through many steps tend to vanish, so plain RNNs forget early words.

LSTM. The LSTM adds a cell state, a memory highway running along the top, and three gates. The forget gate decides what to erase. The input gate decides what new information to store. The output gate decides what to reveal. This lets LSTMs remember much longer.

RNNs today. Since 2017, Transformers have replaced RNNs for most language tasks, because attention looks at all words at once and trains in parallel. But LSTMs remain useful for small, streaming and on-device sequence problems.

Recap. To recap. RNNs carry a hidden state through time with shared weights. Plain RNNs forget over long ranges. LSTMs use gates and a cell state to remember much longer.