Attention and Transformers
Self-attention lets every word look at every other word. See how “it” finds “animal”, and how attention matrices power the Transformer.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. The Transformer, introduced in 2017, is the architecture behind modern language models. Its key idea is attention: letting each word decide which other words matter to it.
Self-attention. Read this sentence. The animal did not cross the street because it was tired. What does it refer to? Self attention lets the word it look at every other word and assign weights. Here most of its attention goes to animal, which is exactly right.
Query, key, value. Attention works with three vectors per word. A query, what am I looking for. A key, what do I contain. Comparing a query with all keys gives weights that add up to one. The word then becomes a weighted blend of all the words’ values.
Attention matrix. Doing this for every word at once gives an attention matrix. Each row shows how much one word attends to every other word. Because every word can look at every other word directly, transformers capture long-range relationships and train efficiently in parallel.
Transformer blocks. A Transformer stacks many identical blocks. Each has multi-head attention, so it can track several kinds of relationships at once, followed by a small feed-forward network, with residual connections to keep training stable.
Recap. To recap. Attention lets every token weigh every other token, using queries, keys and values. Transformers stack attention blocks, and they power today’s language, vision and speech models.