AI in Motion

Natural Language ProcessingIntermediate1:26 video6 chapters

Word2vec and Word Embeddings — lecture notes

You shall know a word by the company it keeps. See how skip-gram turns context windows into training pairs — and meaning into vectors.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Word2vec and Word Embeddings

In 2013, word2vec showed that a simple neural network could learn the meaning of words just by reading lots of text. The key idea: a word is known by the company it keeps.

0:142. Context windows

Context windows — Word2vec and Word Embeddings

Slide a window across the text. For each centre word, every neighbour within two positions becomes a training pair: quick predicts the, brown and fox. Millions of such pairs teach the network which words appear in similar contexts.

0:313. How it trains

How it trains — Word2vec and Word Embeddings

For each pair, look up the centre word’s vector, predict the context word, and nudge the vectors so real neighbours score higher than random words. That trick is called negative sampling.

0:444. Analogies

Analogies — Word2vec and Word Embeddings

After training, similar words sit close together, and directions carry meaning. The step from man to woman is parallel to the step from king to queen, so king minus man plus woman lands near queen.

1:005. Static vs contextual

Static vs contextual — Word2vec and Word Embeddings

Word2vec gives each word a single vector, so bank means the same in river bank and bank loan. Modern models like BERT produce contextual embeddings, a different vector each time, depending on the sentence.

1:156. Recap

Recap — Word2vec and Word Embeddings

To recap. Word2vec learns vectors by predicting neighbours. Similar words end up close together, directions capture relationships, and contextual models took the idea further.

Key takeaways

  • Word2vec (2013) learns word vectors by predicting context words.
  • Skip-gram creates (centre, context) pairs from a sliding window.
  • Negative sampling contrasts real context words with random ones.
  • Static embeddings give one vector per word; contextual models give one per occurrence.

Check yourself

  1. With window size 2, which words are context for “brown” in “the quick brown fox jumps”?
    Show answer

    the, quick, fox, jumps — Two words on each side.

  2. What does negative sampling do?
    Show answer

    Contrasts real context words with random words — It makes training efficient by scoring a few random negatives.

  3. How do contextual embeddings differ from word2vec?
    Show answer

    The same word gets different vectors in different sentences — Context changes the representation.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/word2vec-and-word-embeddings.html