Recurrent Networks and LSTMs — lecture notes
Networks with memory: an RNN reads a sentence word by word, carrying a hidden state; an LSTM adds gates to remember for longer.
0:001. Introduction

Language, music and sensor readings are sequences, where order matters. Recurrent neural networks process sequences one step at a time, carrying a memory forward.
0:112. An RNN unrolled

Here an RNN reads a sentence word by word. At each step it combines the current word with its hidden state, the memory from previous words, and produces a new hidden state. The amber arrow carries that memory forward through time.
0:283. The maths

The new hidden state combines the previous hidden state and the current input, using the same weights at every step. But long sentences cause trouble: gradients passed back through many steps tend to vanish, so plain RNNs forget early words.
0:454. LSTM

The LSTM adds a cell state, a memory highway running along the top, and three gates. The forget gate decides what to erase. The input gate decides what new information to store. The output gate decides what to reveal. This lets LSTMs remember much longer.
1:055. RNNs today

Since 2017, Transformers have replaced RNNs for most language tasks, because attention looks at all words at once and trains in parallel. But LSTMs remain useful for small, streaming and on-device sequence problems.
1:196. Recap

To recap. RNNs carry a hidden state through time with shared weights. Plain RNNs forget over long ranges. LSTMs use gates and a cell state to remember much longer.
Key takeaways
- RNNs process sequences step by step, carrying a hidden state.
- The same weights are reused at every time step.
- LSTMs add a cell state and forget, input and output gates.
- Transformers now dominate language tasks, but LSTMs remain useful for streaming data.
Check yourself
- What carries information from one time step to the next in an RNN?
Show answer
The hidden state — The hidden state is the network’s memory.
- Which LSTM gate decides what to erase from memory?
Show answer
Forget gate — The forget gate scales down parts of the cell state.
- Why did Transformers replace RNNs for most language tasks?
Show answer
Attention sees all tokens at once and trains in parallel — Parallel attention scales far better on modern hardware.
Go deeper
- Recurrent Neural Networks: Modelling Sequences · The AI Lecture Hall
- Long Short-Term Memory (LSTM): Gated Memory Explained · The AI Lecture Hall
- Gated Recurrent Units (GRU) and Choosing a Recurrent Cell · The AI Lecture Hall
- Sequence-to-Sequence Models and the Encoder–Decoder Framework · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/recurrent-networks-and-lstm.html