Bag of Words and TF-IDF — lecture notes
The classic way to turn documents into numbers: count words, then weigh them by how rare they are. Computed live on three tiny documents.
0:001. Introduction

Long before neural networks, search engines and spam filters turned documents into numbers by counting words. The idea is simple and still useful today.
0:112. Count, then weigh

A bag of words counts how often each word appears in each document. Word order is thrown away. But common words like the dominate the counts. TF-IDF fixes that. Words that appear in every document get an inverse document frequency of zero, so the has no weight. Rare, distinctive words like mat, log and chased get the highest weights.
0:353. The formula

TF-IDF multiplies term frequency, how often a word appears in this document, by inverse document frequency, the log of the number of documents divided by how many contain the word. A word found everywhere scores zero.
0:504. Strengths and limits

Bag of words is fast, simple and still strong for search and spam filtering. But it ignores word order, so dog bites man equals man bites dog, and it has no idea that car and automobile mean the same thing. Embeddings solve that.
1:095. Recap

To recap. Bag of words counts words. TF-IDF weighs them by how distinctive they are. It is a great baseline, but it does not understand meaning.
Key takeaways
- Bag of words represents a document by word counts, ignoring order.
- TF-IDF = term frequency × log(N / document frequency).
- Words in every document (like “the”) get zero weight; rare words get high weight.
- It is a strong, fast baseline but has no notion of meaning or order.
Check yourself
- What is the IDF of a word that appears in every document?
Show answer
0 — log(N / N) = log(1) = 0.
- What information does bag of words throw away?
Show answer
Word order — Only counts remain.
- Which words get the highest TF-IDF weight?
Show answer
Frequent in this document but rare elsewhere — High term frequency times high inverse document frequency.
Go deeper
- Bag of Words and TF-IDF: Classical Text Representation · The AI Lecture Hall
- Text Classification: From Linear Models to Fine-Tuned Transformers · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/bag-of-words-and-tf-idf.html