AI in Motion

Machine LearningBeginner1:21 video6 chapters

Linear Regression, Visually — lecture notes

Fit a straight line to data with gradient descent. Watch the residuals shrink and the mean squared error fall from 9.45 to 0.55.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Linear Regression, Visually

Does studying longer lead to higher scores? Linear regression answers questions like this by fitting a straight line through data.

0:092. The model

The model — Linear Regression, Visually

The model is a line: y hat equals w times x plus b. W is the slope, how steep the line is. B is the intercept, where it crosses the axis. Training means finding the best w and b.

0:263. Fitting the line

Fitting the line — Linear Regression, Visually

The red lines are residuals, the errors between each real score and the prediction. Gradient descent nudges w and b to shrink them. The mean squared error starts at 9.45 and falls to about 0.55 after 400 small steps.

0:434. Mean squared error

Mean squared error — Linear Regression, Visually

We measure the fit with the mean squared error. Take each residual, square it so it is positive and big errors count more, then average them. Lower is better.

0:565. Strengths and limits

Strengths and limits — Linear Regression, Visually

Linear regression is simple, fast and easy to explain. But real relationships are often curved, and a few outliers can drag the line off course. Then we add features or use more flexible models.

1:116. Recap

Recap — Linear Regression, Visually

To recap. The model is a line. Residuals are its errors. Mean squared error measures them. And gradient descent finds the best line.

Key takeaways

  • Linear regression fits ŷ = w·x + b.
  • Residuals are the differences between predictions and real values.
  • Mean squared error averages squared residuals; in the animation it falls from 9.45 to 0.55.
  • It is interpretable but limited to straight-line relationships.

Check yourself

  1. In ŷ = w·x + b, what does w control?
    Show answer

    The slope of the line — w is the weight or slope.

  2. Why are errors squared in MSE?
    Show answer

    To make them positive and punish large errors more — Squaring removes signs and emphasises big mistakes.

  3. What does gradient descent adjust during training?
    Show answer

    The parameters w and b — Training updates the model’s parameters.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/linear-regression.html