k-Means Clustering, Step by Step — lecture notes
Unsupervised learning in action: assign points to the nearest centroid, move each centroid to its points’ average, repeat.
0:001. Introduction

You have lots of data and no labels. Are there natural groups hiding in it? k means is the most popular way to find out.
0:112. Watching k-means

We ask for three clusters and drop three centroids, the crosses, at random positions. Step one: every point joins its nearest centroid. Step two: every centroid moves to the average of its points. Repeat. The within-cluster distance on the right drops every round, until the centroids stop moving.
0:313. The algorithm

Choose k. Place the centroids. Assign every point to its nearest centroid. Move each centroid to the mean of its points. Repeat until nothing changes. That is the whole algorithm.
0:454. Pitfalls

You must choose k yourself. The elbow method and silhouette score help. Run it several times, because the start matters. Scale your features. And remember, k means assumes round clusters of similar size.
0:595. Recap

To recap. k means needs no labels. It alternates between assigning points and updating centroids until it converges. It is used for customer segments, image compression and exploring data.
Key takeaways
- k-means alternates between assigning points and moving centroids.
- The within-cluster distance never increases from one round to the next.
- You choose k — the elbow method and silhouette score help.
- Results depend on initial centroids and feature scaling.
Check yourself
- In the update step, each centroid moves to…
Show answer
The mean of its assigned points — Centroids become the average of their cluster members.
- Does k-means need labelled data?
Show answer
No — it is unsupervised — It groups unlabelled data.
- Why run k-means several times?
Show answer
Different starting centroids can give different results — Initialisation affects which solution it converges to.
Go deeper
- k-Means Clustering: Algorithm, Objective and Pitfalls · The AI Lecture Hall
- Hierarchical Clustering and Dendrograms · The AI Lecture Hall
- DBSCAN and Density-Based Clustering · The AI Lecture Hall
- Gaussian Mixture Models and the EM Algorithm · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/k-means-clustering.html