AI in Motion

Artificial IntelligenceIntermediate1:35 video6 chapters

Attention and Transformers — lecture notes

Self-attention lets every word look at every other word. See how “it” finds “animal”, and how attention matrices power the Transformer.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Attention and Transformers

The Transformer, introduced in 2017, is the architecture behind modern language models. Its key idea is attention: letting each word decide which other words matter to it.

0:122. Self-attention

Self-attention — Attention and Transformers

Read this sentence. The animal did not cross the street because it was tired. What does it refer to? Self attention lets the word it look at every other word and assign weights. Here most of its attention goes to animal, which is exactly right.

0:313. Query, key, value

Query, key, value — Attention and Transformers

Attention works with three vectors per word. A query, what am I looking for. A key, what do I contain. Comparing a query with all keys gives weights that add up to one. The word then becomes a weighted blend of all the words’ values.

0:504. Attention matrix

Attention matrix — Attention and Transformers

Doing this for every word at once gives an attention matrix. Each row shows how much one word attends to every other word. Because every word can look at every other word directly, transformers capture long-range relationships and train efficiently in parallel.

1:085. Transformer blocks

Transformer blocks — Attention and Transformers

A Transformer stacks many identical blocks. Each has multi-head attention, so it can track several kinds of relationships at once, followed by a small feed-forward network, with residual connections to keep training stable.

1:236. Recap

Recap — Attention and Transformers

To recap. Attention lets every token weigh every other token, using queries, keys and values. Transformers stack attention blocks, and they power today’s language, vision and speech models.

Key takeaways

  • Self-attention lets each token weigh the relevance of every other token.
  • Queries are compared with keys to produce weights; values are blended with those weights.
  • Each row of an attention matrix sums to 1.
  • Transformers stack multi-head attention and feed-forward layers.

Check yourself

  1. In the example sentence, which word did “it” attend to most?
    Show answer

    animal — “It” refers to the animal, and the attention weights reflect that.

  2. What do the attention weights in one row add up to?
    Show answer

    1 — They are produced by a softmax, so they sum to 1.

  3. Why does multi-head attention use several heads?
    Show answer

    To track different kinds of relationships at the same time — Each head can learn a different attention pattern.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/attention-and-transformers.html