AI in Motion

N-gram Language Models

Natural Language ProcessingBeginner1:276 chapters

Predict the next word by counting word pairs. A bigram model built from a tiny corpus — the ancestor of today’s LLMs.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 In the corpus, “i” is followed by “like” 2 times and “drink” once. What is P(like | i)?
Q2 What problem does smoothing solve?
Q3 How much context does a trigram model use?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. A language model predicts the next word. Today that is done by giant neural networks, but the idea started with simple counting.

A bigram model. Here is a tiny corpus. A bigram model looks at pairs of words. After the word i, the corpus continues with like twice and drink once, so the probability of like is two thirds. After like, green and black are equally likely. After green, it is always tea.

The formula. The probability of the next word, given the previous one, is the count of the pair divided by the count of the previous word. Like after i: two divided by three.

Limitations. Counting has problems. Most word combinations never appear, so they get zero probability. Smoothing techniques give unseen pairs a little probability. And n-grams have short memory: a trigram only sees two words back.

From counts to neural. Neural language models replace counts with learned vectors, generalise to phrases they have never seen, and use attention to look back across thousands of tokens. But the goal is the same: predict the next word.

Recap. To recap. A language model predicts the next word. N-grams do it by counting. Smoothing handles unseen pairs. And modern neural models pursue the same goal with far more context.