Positional Encoding: Teaching Transformers Word Order — lecture notes
Attention ignores order, so Transformers add position information. See the sine-and-cosine pattern that gives every position a unique fingerprint.
0:001. Introduction

Attention compares every word with every other word, but on its own it has no idea which word came first. Dog bites man and man bites dog would look the same. Positional encodings fix that.
0:152. Waves of different speeds

The original Transformer adds a pattern of sine and cosine waves to each word’s embedding. Early dimensions oscillate quickly, later ones slowly, like the hands of a clock. Together they give every position a unique fingerprint, shown as one row of this heat map.
0:343. The formula

Each pair of dimensions uses a sine and cosine at its own frequency, from fast to very slow. A neat property: moving a fixed number of positions corresponds to a rotation, so relative positions are easy for attention to detect.
0:514. Modern variants

Absolute encodings, sinusoidal or learned, give each position its own vector. Many modern language models use relative schemes instead, such as rotary embeddings, RoPE, or ALiBi, which often cope better with longer texts.
1:065. Recap

To recap. Attention is blind to order, so we add positional information. Sinusoidal encodings use waves of many speeds, and modern models often use relative methods like RoPE.
Key takeaways
- Self-attention alone does not know the order of tokens.
- The original Transformer adds sinusoidal positional encodings to embeddings.
- Different dimensions use different frequencies, giving each position a unique pattern.
- Many modern LLMs use relative schemes such as RoPE or ALiBi.
Check yourself
- Why do Transformers need positional encodings?
Show answer
Attention by itself ignores word order — Without them, permuted sentences look identical.
- In sinusoidal encodings, later dimensions…
Show answer
Oscillate more slowly — Frequency decreases with the dimension index.
- Which is a relative position method used by many modern LLMs?
Show answer
RoPE — Rotary position embeddings rotate queries and keys.
Go deeper
- Positional Encodings: Sinusoidal, Learned, RoPE and ALiBi · The AI Lecture Hall
- The Transformer Architecture Explained, Block by Block · The AI Lecture Hall
- Efficient Transformers: Sparse Attention, Linear Attention and FlashAttention · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/positional-encoding.html