AI in Motion

Machine LearningIntermediate1:18 video5 chapters

Principal Component Analysis (PCA) — lecture notes

Find the direction in which data varies most, then project onto it. Watch two dimensions become one while keeping 93.5% of the variance.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Principal Component Analysis (PCA)

Real datasets can have hundreds of features. Principal component analysis, or PCA, finds a smaller set of directions that capture most of the information.

0:112. Finding the first component

Finding the first component — Principal Component Analysis (PCA)

Here is a cloud of points that is stretched in one direction. We rotate a line and measure how much the data spreads along it. The spread, or variance, is highest along one direction. That is the first principal component. Projecting onto it turns two dimensions into one while keeping 93.5 percent of the variance.

0:343. The recipe

The recipe — Principal Component Analysis (PCA)

The recipe: centre the data, compute the covariance matrix, and find its eigenvectors. They point along the directions of greatest variance. Keep the top few and project the data onto them.

0:474. Uses

Uses — Principal Component Analysis (PCA)

PCA is used to visualise high-dimensional data in two dimensions, compress data, remove noise and speed up other models. But the new components are mixtures of features, so they can be harder to interpret.

1:025. Recap

Recap — Principal Component Analysis (PCA)

To recap. Principal components point where the data varies most. They come from eigenvectors of the covariance matrix. Projecting onto the top few reduces dimensions, and you should always check how much variance you keep.

Key takeaways

  • PCA finds orthogonal directions of maximum variance.
  • They are the eigenvectors of the data’s covariance matrix.
  • In the animation, one component keeps 93.5% of the variance.
  • PCA helps with visualisation, compression and noise reduction.

Check yourself

  1. The first principal component is the direction of…
    Show answer

    Maximum variance — PC1 captures the largest spread in the data.

  2. What should you do before PCA?
    Show answer

    Centre (and usually scale) the data — PCA works on centred data; scaling prevents one feature dominating.

  3. What is a downside of PCA?
    Show answer

    Components mix features, so they can be hard to interpret — Each component is a combination of original features.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/principal-component-analysis.html