Natural Language Processing
Tokenization, TF-IDF, n-grams, word2vec, classification, NER, translation, positional encoding, BERT vs GPT, semantic search and speech.
14 animated lectures · 38 minutes
▶ Start with lecture 1What Is Natural Language Processing?
How computers read, understand and generate human language — the tasks, the pipeline and how the field moved from rules to neural networks.
Tokenization and Byte-Pair Encoding
Words, characters or sub-words? See three ways to tokenise text, then watch byte-pair encoding learn sub-words from a tiny corpus.
Bag of Words and TF-IDF
The classic way to turn documents into numbers: count words, then weigh them by how rare they are. Computed live on three tiny documents.
N-gram Language Models
Predict the next word by counting word pairs. A bigram model built from a tiny corpus — the ancestor of today’s LLMs.
Word2vec and Word Embeddings
You shall know a word by the company it keeps. See how skip-gram turns context windows into training pairs — and meaning into vectors.
Sentiment Analysis and Text Classification
Is this review positive or negative? See how word evidence adds up, why negation is tricky, and how modern classifiers are built.
Named Entity Recognition
Find people, organisations, places and dates in text and tag every token with BIO labels.
Sequence-to-Sequence Models and Translation
An encoder reads a sentence, a decoder writes the translation — and attention lets it look back at exactly the right words.
Positional Encoding: Teaching Transformers Word Order
Attention ignores order, so Transformers add position information. See the sine-and-cosine pattern that gives every position a unique fingerprint.
BERT and GPT: Two Kinds of Language Models
BERT reads in both directions to understand; GPT reads left to right to generate. See their training games and attention masks.
Semantic Search with Embeddings
Search by meaning, not matching words. Embed documents and queries, then find the nearest neighbours.
Speech Recognition: From Sound to Text
How a voice becomes words: waveforms, spectrograms and neural models that turn frequency patterns into text.
Word Embeddings: A Deep Dive
How words became vectors: one-hot and TF-IDF, the distributional hypothesis, word2vec skip-gram with negative sampling, analogies, GloVe and fastText, contextual embeddings from BERT, sentence embeddings for search, and bias.
Text Classification from Start to Finish
Build a real text classifier: framing and labelling, tokenisation and TF-IDF features, Naive Bayes and logistic-regression baselines, cross-validation and metrics, imbalance, fine-tuning transformers, zero-shot LLMs, error analysis and monitoring.