AI in Motion

Natural Language ProcessingIntermediate1:18 video5 chapters

Positional Encoding: Teaching Transformers Word Order — lecture notes

Attention ignores order, so Transformers add position information. See the sine-and-cosine pattern that gives every position a unique fingerprint.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Positional Encoding: Teaching Transformers Word Order

Attention compares every word with every other word, but on its own it has no idea which word came first. Dog bites man and man bites dog would look the same. Positional encodings fix that.

0:152. Waves of different speeds

Waves of different speeds — Positional Encoding: Teaching Transformers Word Order

The original Transformer adds a pattern of sine and cosine waves to each word’s embedding. Early dimensions oscillate quickly, later ones slowly, like the hands of a clock. Together they give every position a unique fingerprint, shown as one row of this heat map.

0:343. The formula

The formula — Positional Encoding: Teaching Transformers Word Order

Each pair of dimensions uses a sine and cosine at its own frequency, from fast to very slow. A neat property: moving a fixed number of positions corresponds to a rotation, so relative positions are easy for attention to detect.

0:514. Modern variants

Modern variants — Positional Encoding: Teaching Transformers Word Order

Absolute encodings, sinusoidal or learned, give each position its own vector. Many modern language models use relative schemes instead, such as rotary embeddings, RoPE, or ALiBi, which often cope better with longer texts.

1:065. Recap

Recap — Positional Encoding: Teaching Transformers Word Order

To recap. Attention is blind to order, so we add positional information. Sinusoidal encodings use waves of many speeds, and modern models often use relative methods like RoPE.

Key takeaways

  • Self-attention alone does not know the order of tokens.
  • The original Transformer adds sinusoidal positional encodings to embeddings.
  • Different dimensions use different frequencies, giving each position a unique pattern.
  • Many modern LLMs use relative schemes such as RoPE or ALiBi.

Check yourself

  1. Why do Transformers need positional encodings?
    Show answer

    Attention by itself ignores word order — Without them, permuted sentences look identical.

  2. In sinusoidal encodings, later dimensions…
    Show answer

    Oscillate more slowly — Frequency decreases with the dimension index.

  3. Which is a relative position method used by many modern LLMs?
    Show answer

    RoPE — Rotary position embeddings rotate queries and keys.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/positional-encoding.html