AI in Motion

Machine LearningBeginner1:12 video5 chapters

k-Means Clustering, Step by Step — lecture notes

Unsupervised learning in action: assign points to the nearest centroid, move each centroid to its points’ average, repeat.

▶ Watch the animated lecture

0:001. Introduction

Introduction — k-Means Clustering, Step by Step

You have lots of data and no labels. Are there natural groups hiding in it? k means is the most popular way to find out.

0:112. Watching k-means

Watching k-means — k-Means Clustering, Step by Step

We ask for three clusters and drop three centroids, the crosses, at random positions. Step one: every point joins its nearest centroid. Step two: every centroid moves to the average of its points. Repeat. The within-cluster distance on the right drops every round, until the centroids stop moving.

0:313. The algorithm

The algorithm — k-Means Clustering, Step by Step

Choose k. Place the centroids. Assign every point to its nearest centroid. Move each centroid to the mean of its points. Repeat until nothing changes. That is the whole algorithm.

0:454. Pitfalls

Pitfalls — k-Means Clustering, Step by Step

You must choose k yourself. The elbow method and silhouette score help. Run it several times, because the start matters. Scale your features. And remember, k means assumes round clusters of similar size.

0:595. Recap

Recap — k-Means Clustering, Step by Step

To recap. k means needs no labels. It alternates between assigning points and updating centroids until it converges. It is used for customer segments, image compression and exploring data.

Key takeaways

  • k-means alternates between assigning points and moving centroids.
  • The within-cluster distance never increases from one round to the next.
  • You choose k — the elbow method and silhouette score help.
  • Results depend on initial centroids and feature scaling.

Check yourself

  1. In the update step, each centroid moves to…
    Show answer

    The mean of its assigned points — Centroids become the average of their cluster members.

  2. Does k-means need labelled data?
    Show answer

    No — it is unsupervised — It groups unlabelled data.

  3. Why run k-means several times?
    Show answer

    Different starting centroids can give different results — Initialisation affects which solution it converges to.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/k-means-clustering.html