AI in Motion

Positional Encoding: Teaching Transformers Word Order

Natural Language ProcessingIntermediate1:185 chapters

Attention ignores order, so Transformers add position information. See the sine-and-cosine pattern that gives every position a unique fingerprint.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 Why do Transformers need positional encodings?
Q2 In sinusoidal encodings, later dimensions…
Q3 Which is a relative position method used by many modern LLMs?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. Attention compares every word with every other word, but on its own it has no idea which word came first. Dog bites man and man bites dog would look the same. Positional encodings fix that.

Waves of different speeds. The original Transformer adds a pattern of sine and cosine waves to each word’s embedding. Early dimensions oscillate quickly, later ones slowly, like the hands of a clock. Together they give every position a unique fingerprint, shown as one row of this heat map.

The formula. Each pair of dimensions uses a sine and cosine at its own frequency, from fast to very slow. A neat property: moving a fixed number of positions corresponds to a rotation, so relative positions are easy for attention to detect.

Modern variants. Absolute encodings, sinusoidal or learned, give each position its own vector. Many modern language models use relative schemes instead, such as rotary embeddings, RoPE, or ALiBi, which often cope better with longer texts.

Recap. To recap. Attention is blind to order, so we add positional information. Sinusoidal encodings use waves of many speeds, and modern models often use relative methods like RoPE.