AI in Motion

Natural Language ProcessingIntermediate1:24 video6 chapters

Sequence-to-Sequence Models and Translation — lecture notes

An encoder reads a sentence, a decoder writes the translation — and attention lets it look back at exactly the right words.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Sequence-to-Sequence Models and Translation

Translation turns one sequence into another, of a different length and word order. Sequence to sequence models, introduced around 2014, made neural translation practical.

0:112. Encoder and decoder

Encoder and decoder — Sequence-to-Sequence Models and Translation

The encoder reads the English sentence word by word and compresses it into a context vector. The decoder then writes the French translation one word at a time: le chat noir dort. Notice that the order changes: black cat becomes chat noir.

0:293. The bottleneck

The bottleneck — Sequence-to-Sequence Models and Translation

But squeezing a whole sentence into one fixed vector is a bottleneck. For long sentences, information gets lost, and translation quality drops.

0:394. Adding attention

Adding attention — Sequence-to-Sequence Models and Translation

Attention fixes the bottleneck. At every step, the decoder looks back at all the encoder states and decides which to focus on. When writing chat, it attends to cat. When writing noir, it attends to black. The word order problem solves itself.

0:575. What came next

What came next — Sequence-to-Sequence Models and Translation

Attention was so useful that the 2017 Transformer kept only attention and dropped the recurrent networks. Encoder decoder Transformers such as T5 are still used for translation and summarisation.

1:106. Recap

Recap — Sequence-to-Sequence Models and Translation

To recap. The encoder reads, the decoder writes. A single context vector is a bottleneck. Attention lets the decoder look back at every word, and that idea led straight to the Transformer.

Key takeaways

  • Seq2seq models use an encoder to read and a decoder to generate.
  • A single fixed context vector limits quality on long sentences.
  • Attention lets the decoder focus on relevant source words at every step.
  • The Transformer (2017) built an entire architecture from attention.

Check yourself

  1. What does the encoder do?
    Show answer

    Reads the source sentence into representations — The encoder processes the input sequence.

  2. What problem does attention solve in seq2seq models?
    Show answer

    The fixed-size bottleneck vector — The decoder can look at all encoder states directly.

  3. When writing “noir”, the decoder should attend mostly to…
    Show answer

    black — “Noir” translates “black”.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/sequence-to-sequence-translation.html