AI in Motion

Natural Language ProcessingBeginner1:40 video6 chapters

Tokenization and Byte-Pair Encoding — lecture notes

Words, characters or sub-words? See three ways to tokenise text, then watch byte-pair encoding learn sub-words from a tiny corpus.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Tokenization and Byte-Pair Encoding

Before a model can read a sentence, the sentence must be cut into pieces called tokens. How we cut it matters more than you might think.

0:112. Three strategies

Three strategies — Tokenization and Byte-Pair Encoding

Word-level tokenization gives short sequences, but needs a huge vocabulary and cannot handle new words. Character-level needs only a tiny vocabulary, but sequences become very long. Sub-word tokenization sits in between: common words stay whole, rare words split into meaningful pieces.

0:293. Byte-pair encoding

Byte-pair encoding — Tokenization and Byte-Pair Encoding

Byte-pair encoding learns sub-words from data. Start with single characters. Count every adjacent pair, weighted by word frequency, and merge the most frequent one. Here e and s appear together nine times, so they merge first, then e s with t, then est with the end of a word. Repeat, and useful pieces like est and low emerge.

0:534. Why sub-words

Why sub-words — Tokenization and Byte-Pair Encoding

Sub-words mean no word is ever unknown, because anything can be spelled from pieces. Related words share pieces, and the vocabulary stays manageable. GPT models use byte-pair encoding, and BERT uses a close cousin called WordPiece.

1:095. In practice

In practice — Tokenization and Byte-Pair Encoding

In English, a token is roughly three quarters of a word on average. Many other languages need more tokens for the same meaning. Model limits and prices are counted in tokens, and tokenization even explains why language models struggle to count letters in a word.

1:286. Recap

Recap — Tokenization and Byte-Pair Encoding

To recap. Tokens are what models read. Sub-word tokenization is the modern standard, and byte-pair encoding builds it by repeatedly merging the most frequent pair of symbols.

Key takeaways

  • Tokenization splits text into units (tokens) that a model processes.
  • Sub-word tokenization balances vocabulary size and sequence length.
  • BPE repeatedly merges the most frequent adjacent pair; in the demo e+s (9 times) merges first.
  • In English, a token is about ¾ of a word on average.

Check yourself

  1. What does each BPE step do?
    Show answer

    Merges the most frequent adjacent pair of symbols — BPE builds sub-words by merging the most common pair.

  2. A key advantage of sub-word tokenization is…
    Show answer

    It can represent any word, even unseen ones — Unknown words can be spelled from known pieces.

  3. Why can character-level models be slow?
    Show answer

    Sequences become very long — Every character is a separate token.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/tokenization-and-subwords.html