AI in Motion

k-Means Clustering, Step by Step

Machine LearningBeginner1:125 chapters

Unsupervised learning in action: assign points to the nearest centroid, move each centroid to its points’ average, repeat.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 In the update step, each centroid moves to…
Q2 Does k-means need labelled data?
Q3 Why run k-means several times?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. You have lots of data and no labels. Are there natural groups hiding in it? k means is the most popular way to find out.

Watching k-means. We ask for three clusters and drop three centroids, the crosses, at random positions. Step one: every point joins its nearest centroid. Step two: every centroid moves to the average of its points. Repeat. The within-cluster distance on the right drops every round, until the centroids stop moving.

The algorithm. Choose k. Place the centroids. Assign every point to its nearest centroid. Move each centroid to the mean of its points. Repeat until nothing changes. That is the whole algorithm.

Pitfalls. You must choose k yourself. The elbow method and silhouette score help. Run it several times, because the start matters. Scale your features. And remember, k means assumes round clusters of similar size.

Recap. To recap. k means needs no labels. It alternates between assigning points and updating centroids until it converges. It is used for customer segments, image compression and exploring data.