Positional Encoding: Teaching Transformers Word Order
Attention ignores order, so Transformers add position information. See the sine-and-cosine pattern that gives every position a unique fingerprint.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. Attention compares every word with every other word, but on its own it has no idea which word came first. Dog bites man and man bites dog would look the same. Positional encodings fix that.
Waves of different speeds. The original Transformer adds a pattern of sine and cosine waves to each word’s embedding. Early dimensions oscillate quickly, later ones slowly, like the hands of a clock. Together they give every position a unique fingerprint, shown as one row of this heat map.
The formula. Each pair of dimensions uses a sine and cosine at its own frequency, from fast to very slow. A neat property: moving a fixed number of positions corresponds to a rotation, so relative positions are easy for attention to detect.
Modern variants. Absolute encodings, sinusoidal or learned, give each position its own vector. Many modern language models use relative schemes instead, such as rotary embeddings, RoPE, or ALiBi, which often cope better with longer texts.
Recap. To recap. Attention is blind to order, so we add positional information. Sinusoidal encodings use waves of many speeds, and modern models often use relative methods like RoPE.