k-Means Clustering, Step by Step
Unsupervised learning in action: assign points to the nearest centroid, move each centroid to its points’ average, repeat.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. You have lots of data and no labels. Are there natural groups hiding in it? k means is the most popular way to find out.
Watching k-means. We ask for three clusters and drop three centroids, the crosses, at random positions. Step one: every point joins its nearest centroid. Step two: every centroid moves to the average of its points. Repeat. The within-cluster distance on the right drops every round, until the centroids stop moving.
The algorithm. Choose k. Place the centroids. Assign every point to its nearest centroid. Move each centroid to the mean of its points. Repeat until nothing changes. That is the whole algorithm.
Pitfalls. You must choose k yourself. The elbow method and silhouette score help. Run it several times, because the start matters. Scale your features. And remember, k means assumes round clusters of similar size.
Recap. To recap. k means needs no labels. It alternates between assigning points and updating centroids until it converges. It is used for customer segments, image compression and exploring data.