Sequence-to-Sequence Models and Translation
An encoder reads a sentence, a decoder writes the translation — and attention lets it look back at exactly the right words.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. Translation turns one sequence into another, of a different length and word order. Sequence to sequence models, introduced around 2014, made neural translation practical.
Encoder and decoder. The encoder reads the English sentence word by word and compresses it into a context vector. The decoder then writes the French translation one word at a time: le chat noir dort. Notice that the order changes: black cat becomes chat noir.
The bottleneck. But squeezing a whole sentence into one fixed vector is a bottleneck. For long sentences, information gets lost, and translation quality drops.
Adding attention. Attention fixes the bottleneck. At every step, the decoder looks back at all the encoder states and decides which to focus on. When writing chat, it attends to cat. When writing noir, it attends to black. The word order problem solves itself.
What came next. Attention was so useful that the 2017 Transformer kept only attention and dropped the recurrent networks. Encoder decoder Transformers such as T5 are still used for translation and summarisation.
Recap. To recap. The encoder reads, the decoder writes. A single context vector is a bottleneck. Attention lets the decoder look back at every word, and that idea led straight to the Transformer.