Bag of Words and TF-IDF
The classic way to turn documents into numbers: count words, then weigh them by how rare they are. Computed live on three tiny documents.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. Long before neural networks, search engines and spam filters turned documents into numbers by counting words. The idea is simple and still useful today.
Count, then weigh. A bag of words counts how often each word appears in each document. Word order is thrown away. But common words like the dominate the counts. TF-IDF fixes that. Words that appear in every document get an inverse document frequency of zero, so the has no weight. Rare, distinctive words like mat, log and chased get the highest weights.
The formula. TF-IDF multiplies term frequency, how often a word appears in this document, by inverse document frequency, the log of the number of documents divided by how many contain the word. A word found everywhere scores zero.
Strengths and limits. Bag of words is fast, simple and still strong for search and spam filtering. But it ignores word order, so dog bites man equals man bites dog, and it has no idea that car and automobile mean the same thing. Embeddings solve that.
Recap. To recap. Bag of words counts words. TF-IDF weighs them by how distinctive they are. It is a great baseline, but it does not understand meaning.