Overfitting, Underfitting and Bias–Variance — lecture notes
Too simple, just right, too complex: see three models on the same data and why validation error — not training error — tells the truth.
0:001. Introduction

The goal of machine learning is not to fit the training data. It is to do well on new data. Two opposite failures get in the way.
0:122. Three models

The same data, three models. A straight line underfits: training error 1.53. A gentle curve fits well: 0.25. A degree nine polynomial passes through every point: training error zero. But on new validation points, the curve wins with 1.26, while the overfit model is worst at 2.30.
0:323. Bias and variance

Think of darts. High bias means aiming off-centre, like a model that is too simple. High variance means throws scattered everywhere, like a model that changes wildly with small changes in data. We want low bias and low variance.
0:494. Learning curves

Learning curves reveal overfitting during training. Training loss keeps falling, but validation loss stops improving and starts rising. The widening gap means the model is memorising. Early stopping keeps the model from the best point.
1:045. Fixes

To fix underfitting, use a more flexible model, better features, or train longer. To fix overfitting, get more data, simplify or regularise the model, and use tricks like early stopping, dropout and data augmentation.
1:196. Recap

To recap. Underfitting is too simple. Overfitting memorises. Always judge by validation error, and aim for the balance between bias and variance.
Key takeaways
- Underfitting: high error on training and validation data.
- Overfitting: very low training error but high validation error.
- In the animation: validation error 1.87 (line), 1.26 (curve), 2.30 (degree-9 polynomial).
- More data, regularisation and early stopping reduce overfitting.
Check yourself
- A model has 0 training error but high validation error. It is…
Show answer
Overfitting — It has memorised the training data.
- In the darts analogy, high variance looks like…
Show answer
Throws scattered widely — Variance is sensitivity: results scatter.
- Which helps against overfitting?
Show answer
More training data — More data makes memorisation harder and patterns clearer.
Go deeper
- The Bias–Variance Trade-off: Derivation and Intuition · The AI Lecture Hall
- Overfitting and Underfitting: Diagnosis with Learning Curves · The AI Lecture Hall
- Regularisation: Ridge, Lasso and Elastic Net · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/overfitting-and-bias-variance.html