AI in Motion

Natural Language ProcessingIntermediate1:23 video6 chapters

BERT and GPT: Two Kinds of Language Models — lecture notes

BERT reads in both directions to understand; GPT reads left to right to generate. See their training games and attention masks.

▶ Watch the animated lecture

0:001. Introduction

Introduction — BERT and GPT: Two Kinds of Language Models

In 2018, two families of pre-trained Transformer models changed NLP: BERT and GPT. They share building blocks but play different training games.

0:102. BERT fills in blanks

BERT fills in blanks — BERT and GPT: Two Kinds of Language Models

BERT is trained by hiding some words and predicting them. To guess the masked word, it uses context from both sides: the chef cooked a something dinner for us. Reading in both directions makes BERT excellent at understanding text.

0:273. GPT continues text

GPT continues text — BERT and GPT: Two Kinds of Language Models

GPT is trained to predict the next token, reading only from left to right. That makes it a natural writer: give it a beginning and it continues, one token at a time.

0:414. Attention masks

Attention masks — BERT and GPT: Two Kinds of Language Models

The difference shows up in the attention mask. In BERT, every word can attend to every other word. In GPT, a causal mask blocks the future: each word only sees the words before it. Otherwise the model could cheat by looking at the answer.

1:005. Which one?

Which one? — BERT and GPT: Two Kinds of Language Models

BERT-style encoders shine at classification, entity recognition and search embeddings. GPT-style decoders generate text, and scaled up, they became today’s large language models. Encoder decoder models like T5 combine both.

1:136. Recap

Recap — BERT and GPT: Two Kinds of Language Models

To recap. BERT fills in blanks using both directions. GPT predicts the next token using only the past. Encoders understand, decoders generate.

Key takeaways

  • BERT is pre-trained with masked language modelling and reads bidirectionally.
  • GPT is pre-trained to predict the next token, left to right.
  • GPT uses a causal attention mask so tokens cannot see the future.
  • Encoders suit understanding tasks; decoders suit generation.

Check yourself

  1. How is BERT pre-trained?
    Show answer

    Predicting masked words using both sides of context — Masked language modelling is bidirectional.

  2. What does a causal attention mask do?
    Show answer

    Prevents tokens from attending to future tokens — Each position sees only earlier positions.

  3. Which model type is best suited to writing new text?
    Show answer

    GPT-style decoder — Decoders generate token by token.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/bert-and-gpt.html