The Attention Mechanism in Depth — lecture notes
Queries, keys and values; scaled dot-product attention computed by hand; masking; multi-head attention; positional encodings and RoPE; the quadratic cost and how FlashAttention, GQA and sliding windows tame it.
0:001. Introduction

Attention is the single idea that made transformers, and therefore modern language models, possible. The 2017 paper that introduced the transformer was even titled Attention is all you need. In this deep dive we compute attention by hand, then follow it all the way to today’s efficient implementations.
0:202. The problem

Why do we need attention? Because words change meaning with context. Bank means something different in river bank and bank account, and the word it must be linked to whatever it refers to, however far back. A model needs a way to pull information from any other word, directly.
0:413. Where attention began

Attention first appeared in machine translation in 2014. Instead of squeezing a whole sentence into one vector, the decoder was allowed to look back at every source word and focus on the relevant ones while producing each output word. The transformer then used attention everywhere, and dropped recurrence entirely.
1:024. Queries, keys and values

Attention uses three roles. Each token produces a query, saying what it is looking for, a key, advertising what it contains, and a value, the information it will share. Queries are compared with keys, and the values of the best matches are mixed together, like matching a question against book labels and reading the matching books.
1:255. Learned projections

Where do queries, keys and values come from? Each is a learned linear projection of the same token vectors, using three weight matrices. Training decides what a token should look for, what it should advertise, and what it should share. In GPT two small, each projection is a seven hundred and sixty eight square matrix per layer.
1:496. The formula

Here is the whole mechanism in one line. Multiply the queries by the transposed keys to score every pair of tokens. Divide by the square root of the key dimension. Apply a softmax along each row to turn scores into weights. Then use those weights to mix the value vectors.
2:107. Computing it by hand

Let us compute it for one query, the word it. Each score is the query dotted with a key, divided by the square root of three. The softmax turns the scores into weights: the animal receives about thirty nine percent. The output is the weighted sum of the value vectors, dominated by the animal’s value.
2:338. Pause and think

Pause and think. The animal scored point nine four and tired scored point five eight. Why are their weights not simply proportional to those scores? Because softmax exponentiates. Weights are proportional to e to the score, so a difference of point three six becomes a ratio of about one point four.
2:559. Why √d?

Why divide by the square root of the dimension? For random vectors with d components, dot products have a spread of about the square root of d. With sixty four dimensions, raw scores would spread over plus or minus eight, making the softmax almost one hot with vanishing gradients. Scaling restores a healthy range.
3:1810. The attention matrix

Computed for every query at once, attention produces a matrix: one row per token that is looking, one column per token being looked at. Each row is a probability distribution summing to one. Bright cells show where information flows. This entire matrix comes from a single matrix product and a softmax.
3:3911. Pause and think

Pause and think. What does each row of the attention weight matrix always add up to? Exactly one, because each row is a softmax over the keys. It is a probability distribution describing how one token divides its attention among all the tokens it can see.
3:5912. Self vs cross

In self attention, queries, keys and values all come from the same sequence, so a sentence attends to itself. In cross attention, the queries come from one sequence and the keys and values from another. A translation decoder cross attends to the source sentence, and text to image models cross attend to the prompt.
4:2213. Causal masking

Language models need causal masking. Before the softmax, every score for a future position is set to minus infinity, which becomes a weight of exactly zero. Each token can then only gather information from itself and the past, so the model can be trained to predict every next token in parallel without cheating.
4:4414. Pause and think

Pause and think. Why can a causal language model compute the loss for every position of a sentence in one single forward pass? Because the mask guarantees each position sees only earlier tokens, so all the next token predictions can be computed simultaneously without leaking the answers.
5:0415. Multi-head attention

One attention pattern is not enough, so transformers run several heads in parallel. Each head has its own query, key and value projections into a smaller space, and learns its own notion of relevance: previous token, coreference, the first token, or broad context. Their outputs are concatenated and mixed.
5:2516. The multi-head formula

Formally, each head applies attention to its own projections of the input, and the head outputs are concatenated and multiplied by an output matrix, W O. Because each head works in a smaller subspace, multi head attention costs about the same as a single full sized head.
5:4517. Heads in practice

In practice the head dimension is usually sixty four or one hundred and twenty eight. GPT two small uses twelve heads of sixty four dimensions. Llama two seven B uses thirty two heads of one hundred and twenty eight, and GPT three uses ninety six heads of one hundred and twenty eight, in every one of its ninety six layers.
6:0918. Positions, the classic way

Attention by itself is order blind: shuffle the words and the scores would not change. So position must be injected. The original transformer added fixed sinusoidal patterns to the embeddings, with each dimension oscillating at a different frequency, so every position has a unique fingerprint.
6:2819. Rotary position embeddings

Most modern models use rotary position embeddings, RoPE. Instead of adding a pattern, each pair of query and key dimensions is rotated by an angle proportional to the position. The dot product then depends only on the difference in positions: shift both tokens along and the score stays exactly the same.
6:5020. Pause and think

Pause and think. Without any positional information, what would attention make of dog bites man versus man bites dog? They would look identical, because attention treats its input as a set of tokens. Positional encodings are what allow the model to tell who bit whom.
7:0921. Relative positions

RoPE, and alternatives such as ALiBi, which adds a penalty growing with distance, encode relative positions. This helps models generalise across positions, and tricks like rescaling the RoPE frequencies let a model trained on a few thousand tokens be extended to much longer contexts.
7:2822. The quadratic cost

Attention has a price. For n tokens it computes an n by n matrix of scores, so time grows with n squared, and a naive implementation stores the whole matrix. Making the context eight times longer means sixty four times more attention work. This is the main obstacle to very long contexts.
7:5023. Cost of long contexts

Here that cost is laid out. Going from four thousand to thirty two thousand tokens means sixty four times the attention work, and a million tokens would be about sixty thousand times more. Long context models therefore rely on clever algorithms and hardware aware implementations.
8:0924. FlashAttention

FlashAttention, introduced in 2022, computes exactly the same result, but in tiles small enough to fit in the GPU’s fast on chip memory, using a running softmax so the full n by n matrix is never written to slow memory. Less memory traffic makes attention several times faster.
8:3025. Sharing keys and values

Another trick targets inference memory. In multi query attention, all query heads share a single set of keys and values. Grouped query attention is a middle ground with a few shared sets. Llama two seventy B uses sixty four query heads but only eight key value heads, shrinking its memory cache eight fold.
8:5226. Sparse and local attention

Sliding window attention lets each token attend only to a fixed number of recent tokens, making cost grow linearly with length. Because layers are stacked, information can still travel far, one window per layer. Mistral seven B uses a four thousand token window, and other models mix local and global layers.
9:1427. Attention in code

Here is attention in a dozen lines of PyTorch. Compute scaled scores, fill the future positions with minus infinity, apply the softmax, and multiply by the values. In real code you would call the built in fused function, which picks an efficient kernel such as FlashAttention automatically.
9:3428. What heads learn

Interpretability research has found specific circuits inside attention. Induction heads, for example, find where the current token appeared earlier and predict whatever followed it last time: after Mr and Mrs Dursley appears once, the model completes the pattern. These heads appear to underlie much of in context learning.
9:5429. Attention sinks

Another curious finding is the attention sink. Trained models often place a large share of their attention on the first few tokens, using them as a harmless place to park attention when nothing else is relevant. Keeping those tokens in the cache turns out to be essential when streaming very long texts with a sliding window.
10:1830. Attention in context

Let us place attention back into the whole model. In every block, attention is the only step where information moves between positions. The MLP then processes each position separately. Dozens of these blocks, each with many heads, let information hop from word to word until the final position holds what it needs to predict the next token.
10:4231. Attention variants

Here are the variants side by side. Multi head attention gives richer patterns. Causal masking enables next token training. RoPE and ALiBi handle positions. FlashAttention speeds everything up exactly. Grouped query attention shrinks the memory cache, and sliding windows make cost grow linearly with length.
11:0132. Recap

To recap. Every token emits a query, a key and a value. Attention computes the softmax of scaled query key scores and uses it to mix values. A causal mask enables next token training, multiple heads learn different patterns, RoPE adds relative position, and efficient variants tame the quadratic cost.
Key takeaways
- Attention lets every token gather information directly from any other token.
- Attention(Q, K, V) = softmax(QKᵀ / √dₖ) V; each row of weights sums to 1.
- Scaling by √dₖ keeps scores near unit variance so the softmax does not saturate.
- Causal masking sets future scores to −∞, enabling parallel next-token training.
- Multi-head attention runs several heads in smaller subspaces; RoPE encodes relative position by rotation.
- Cost grows as n²; FlashAttention, grouped-query attention and sliding windows make long contexts practical.
Check yourself
- In attention, what are queries compared with?
Show answer
Keys — Scores are query · key dot products.
- Why divide the scores by √dₖ?
Show answer
To keep their scale near 1 so the softmax does not saturate — Dot products of d-dimensional vectors grow like √d.
- How is causal masking implemented?
Show answer
By setting future scores to −∞ before the softmax — e^−∞ = 0, so future weights vanish.
- What does grouped-query attention mainly reduce?
Show answer
The size of the key/value cache during inference — Query heads share fewer key/value heads.
- If the context length doubles, naive attention work grows about…
Show answer
4× — n² scaling.
Go deeper
- Self-Attention in Depth: Intuition, Complexity and Variants · The AI Lecture Hall
- The Attention Mechanism: Learning Where to Look · The AI Lecture Hall
- Positional Encodings: Sinusoidal, Learned, RoPE and ALiBi · The AI Lecture Hall
- Efficient Transformers: Sparse Attention, Linear Attention and FlashAttention · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/attention-mechanism-deep-dive.html