BERT and GPT: Two Kinds of Language Models
BERT reads in both directions to understand; GPT reads left to right to generate. See their training games and attention masks.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. In 2018, two families of pre-trained Transformer models changed NLP: BERT and GPT. They share building blocks but play different training games.
BERT fills in blanks. BERT is trained by hiding some words and predicting them. To guess the masked word, it uses context from both sides: the chef cooked a something dinner for us. Reading in both directions makes BERT excellent at understanding text.
GPT continues text. GPT is trained to predict the next token, reading only from left to right. That makes it a natural writer: give it a beginning and it continues, one token at a time.
Attention masks. The difference shows up in the attention mask. In BERT, every word can attend to every other word. In GPT, a causal mask blocks the future: each word only sees the words before it. Otherwise the model could cheat by looking at the answer.
Which one?. BERT-style encoders shine at classification, entity recognition and search embeddings. GPT-style decoders generate text, and scaled up, they became today’s large language models. Encoder decoder models like T5 combine both.
Recap. To recap. BERT fills in blanks using both directions. GPT predicts the next token using only the past. Encoders understand, decoders generate.