AI in Motion

Sequence-to-Sequence Models and Translation

Natural Language ProcessingIntermediate1:246 chapters

An encoder reads a sentence, a decoder writes the translation — and attention lets it look back at exactly the right words.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 What does the encoder do?
Q2 What problem does attention solve in seq2seq models?
Q3 When writing “noir”, the decoder should attend mostly to…

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. Translation turns one sequence into another, of a different length and word order. Sequence to sequence models, introduced around 2014, made neural translation practical.

Encoder and decoder. The encoder reads the English sentence word by word and compresses it into a context vector. The decoder then writes the French translation one word at a time: le chat noir dort. Notice that the order changes: black cat becomes chat noir.

The bottleneck. But squeezing a whole sentence into one fixed vector is a bottleneck. For long sentences, information gets lost, and translation quality drops.

Adding attention. Attention fixes the bottleneck. At every step, the decoder looks back at all the encoder states and decides which to focus on. When writing chat, it attends to cat. When writing noir, it attends to black. The word order problem solves itself.

What came next. Attention was so useful that the 2017 Transformer kept only attention and dropped the recurrent networks. Encoder decoder Transformers such as T5 are still used for translation and summarisation.

Recap. To recap. The encoder reads, the decoder writes. A single context vector is a bottleneck. Attention lets the decoder look back at every word, and that idea led straight to the Transformer.