Attention and Transformers — lecture notes
Self-attention lets every word look at every other word. See how “it” finds “animal”, and how attention matrices power the Transformer.
0:001. Introduction

The Transformer, introduced in 2017, is the architecture behind modern language models. Its key idea is attention: letting each word decide which other words matter to it.
0:122. Self-attention

Read this sentence. The animal did not cross the street because it was tired. What does it refer to? Self attention lets the word it look at every other word and assign weights. Here most of its attention goes to animal, which is exactly right.
0:313. Query, key, value

Attention works with three vectors per word. A query, what am I looking for. A key, what do I contain. Comparing a query with all keys gives weights that add up to one. The word then becomes a weighted blend of all the words’ values.
0:504. Attention matrix

Doing this for every word at once gives an attention matrix. Each row shows how much one word attends to every other word. Because every word can look at every other word directly, transformers capture long-range relationships and train efficiently in parallel.
1:085. Transformer blocks

A Transformer stacks many identical blocks. Each has multi-head attention, so it can track several kinds of relationships at once, followed by a small feed-forward network, with residual connections to keep training stable.
1:236. Recap

To recap. Attention lets every token weigh every other token, using queries, keys and values. Transformers stack attention blocks, and they power today’s language, vision and speech models.
Key takeaways
- Self-attention lets each token weigh the relevance of every other token.
- Queries are compared with keys to produce weights; values are blended with those weights.
- Each row of an attention matrix sums to 1.
- Transformers stack multi-head attention and feed-forward layers.
Check yourself
- In the example sentence, which word did “it” attend to most?
Show answer
animal — “It” refers to the animal, and the attention weights reflect that.
- What do the attention weights in one row add up to?
Show answer
1 — They are produced by a softmax, so they sum to 1.
- Why does multi-head attention use several heads?
Show answer
To track different kinds of relationships at the same time — Each head can learn a different attention pattern.
Go deeper
- The Attention Mechanism: Learning Where to Look · The AI Lecture Hall
- Self-Attention in Depth: Intuition, Complexity and Variants · The AI Lecture Hall
- The Transformer Architecture Explained, Block by Block · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/attention-and-transformers.html