Tokenization and Byte-Pair Encoding
Words, characters or sub-words? See three ways to tokenise text, then watch byte-pair encoding learn sub-words from a tiny corpus.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. Before a model can read a sentence, the sentence must be cut into pieces called tokens. How we cut it matters more than you might think.
Three strategies. Word-level tokenization gives short sequences, but needs a huge vocabulary and cannot handle new words. Character-level needs only a tiny vocabulary, but sequences become very long. Sub-word tokenization sits in between: common words stay whole, rare words split into meaningful pieces.
Byte-pair encoding. Byte-pair encoding learns sub-words from data. Start with single characters. Count every adjacent pair, weighted by word frequency, and merge the most frequent one. Here e and s appear together nine times, so they merge first, then e s with t, then est with the end of a word. Repeat, and useful pieces like est and low emerge.
Why sub-words. Sub-words mean no word is ever unknown, because anything can be spelled from pieces. Related words share pieces, and the vocabulary stays manageable. GPT models use byte-pair encoding, and BERT uses a close cousin called WordPiece.
In practice. In English, a token is roughly three quarters of a word on average. Many other languages need more tokens for the same meaning. Model limits and prices are counted in tokens, and tokenization even explains why language models struggle to count letters in a word.
Recap. To recap. Tokens are what models read. Sub-word tokenization is the modern standard, and byte-pair encoding builds it by repeatedly merging the most frequent pair of symbols.