AI in Motion

Word2vec and Word Embeddings

Natural Language ProcessingIntermediate1:266 chapters

You shall know a word by the company it keeps. See how skip-gram turns context windows into training pairs — and meaning into vectors.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 With window size 2, which words are context for “brown” in “the quick brown fox jumps”?
Q2 What does negative sampling do?
Q3 How do contextual embeddings differ from word2vec?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. In 2013, word2vec showed that a simple neural network could learn the meaning of words just by reading lots of text. The key idea: a word is known by the company it keeps.

Context windows. Slide a window across the text. For each centre word, every neighbour within two positions becomes a training pair: quick predicts the, brown and fox. Millions of such pairs teach the network which words appear in similar contexts.

How it trains. For each pair, look up the centre word’s vector, predict the context word, and nudge the vectors so real neighbours score higher than random words. That trick is called negative sampling.

Analogies. After training, similar words sit close together, and directions carry meaning. The step from man to woman is parallel to the step from king to queen, so king minus man plus woman lands near queen.

Static vs contextual. Word2vec gives each word a single vector, so bank means the same in river bank and bank loan. Modern models like BERT produce contextual embeddings, a different vector each time, depending on the sentence.

Recap. To recap. Word2vec learns vectors by predicting neighbours. Similar words end up close together, directions capture relationships, and contextual models took the idea further.