AI in Motion

Tokenization and Byte-Pair Encoding

Natural Language ProcessingBeginner1:406 chapters

Words, characters or sub-words? See three ways to tokenise text, then watch byte-pair encoding learn sub-words from a tiny corpus.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 What does each BPE step do?
Q2 A key advantage of sub-word tokenization is…
Q3 Why can character-level models be slow?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. Before a model can read a sentence, the sentence must be cut into pieces called tokens. How we cut it matters more than you might think.

Three strategies. Word-level tokenization gives short sequences, but needs a huge vocabulary and cannot handle new words. Character-level needs only a tiny vocabulary, but sequences become very long. Sub-word tokenization sits in between: common words stay whole, rare words split into meaningful pieces.

Byte-pair encoding. Byte-pair encoding learns sub-words from data. Start with single characters. Count every adjacent pair, weighted by word frequency, and merge the most frequent one. Here e and s appear together nine times, so they merge first, then e s with t, then est with the end of a word. Repeat, and useful pieces like est and low emerge.

Why sub-words. Sub-words mean no word is ever unknown, because anything can be spelled from pieces. Related words share pieces, and the vocabulary stays manageable. GPT models use byte-pair encoding, and BERT uses a close cousin called WordPiece.

In practice. In English, a token is roughly three quarters of a word on average. Many other languages need more tokens for the same meaning. Model limits and prices are counted in tokens, and tokenization even explains why language models struggle to count letters in a word.

Recap. To recap. Tokens are what models read. Sub-word tokenization is the modern standard, and byte-pair encoding builds it by repeatedly merging the most frequent pair of symbols.