BERT and GPT: Two Kinds of Language Models — lecture notes
BERT reads in both directions to understand; GPT reads left to right to generate. See their training games and attention masks.
0:001. Introduction

In 2018, two families of pre-trained Transformer models changed NLP: BERT and GPT. They share building blocks but play different training games.
0:102. BERT fills in blanks

BERT is trained by hiding some words and predicting them. To guess the masked word, it uses context from both sides: the chef cooked a something dinner for us. Reading in both directions makes BERT excellent at understanding text.
0:273. GPT continues text

GPT is trained to predict the next token, reading only from left to right. That makes it a natural writer: give it a beginning and it continues, one token at a time.
0:414. Attention masks

The difference shows up in the attention mask. In BERT, every word can attend to every other word. In GPT, a causal mask blocks the future: each word only sees the words before it. Otherwise the model could cheat by looking at the answer.
1:005. Which one?

BERT-style encoders shine at classification, entity recognition and search embeddings. GPT-style decoders generate text, and scaled up, they became today’s large language models. Encoder decoder models like T5 combine both.
1:136. Recap

To recap. BERT fills in blanks using both directions. GPT predicts the next token using only the past. Encoders understand, decoders generate.
Key takeaways
- BERT is pre-trained with masked language modelling and reads bidirectionally.
- GPT is pre-trained to predict the next token, left to right.
- GPT uses a causal attention mask so tokens cannot see the future.
- Encoders suit understanding tasks; decoders suit generation.
Check yourself
- How is BERT pre-trained?
Show answer
Predicting masked words using both sides of context — Masked language modelling is bidirectional.
- What does a causal attention mask do?
Show answer
Prevents tokens from attending to future tokens — Each position sees only earlier positions.
- Which model type is best suited to writing new text?
Show answer
GPT-style decoder — Decoders generate token by token.
Go deeper
- BERT: Bidirectional Encoder Representations from Transformers · The AI Lecture Hall
- The GPT Family: Autoregressive Language Models from GPT-1 to Today · The AI Lecture Hall
- T5 and BART: Encoder–Decoder Pretraining and Text-to-Text Learning · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/bert-and-gpt.html