Decision Trees and Random Forests — lecture notes
A tree learns yes/no questions that carve the data into pure regions. A forest of trees votes for more robust predictions.
0:001. Introduction

A doctor diagnoses by asking questions. A decision tree learns its own series of yes or no questions from data.
0:092. Growing a tree

Before any split, predicting pink for everything is right 60 percent of the time. The first question, is feature one below 5.5, raises accuracy to 77 percent. A second split makes it 87 percent, and a third makes every region pure: 100 percent on this training data.
0:293. Choosing splits

At each step, the tree tries many possible questions and measures how pure the resulting groups would be, using Gini impurity or entropy. It keeps the best question and repeats on each branch.
0:444. Overfitting

A tree grown too deep memorises noise. It scores perfectly on training data and badly on new data. We control it by limiting depth, requiring enough examples per leaf, or pruning.
0:575. Random forests

A random forest trains many trees, each on a random sample of the data and features. Each tree votes. Here four trees say cat and one says dog, so the forest predicts cat. Averaging many different trees makes predictions far more stable.
1:156. Recap

To recap. Decision trees learn readable questions. Splits aim for purity. Deep trees overfit. And random forests combine many trees into a strong, stable model.
Key takeaways
- Decision trees split data with yes/no questions chosen to maximise purity.
- In the animation, accuracy rises 60% → 77% → 87% → 100% with three splits.
- Deep trees overfit; limit depth or prune.
- Random forests train many randomised trees and let them vote.
Check yourself
- How does a tree choose each question?
Show answer
The split that makes the resulting groups purest — Splits are scored with impurity measures such as Gini or entropy.
- Why can a very deep tree be a problem?
Show answer
It can memorise noise and generalise poorly — Deep trees overfit the training data.
- A random forest makes its final prediction by…
Show answer
Majority vote or averaging across trees — Many trees vote to produce a robust prediction.
Go deeper
- Decision Trees: Splitting Criteria, Pruning and Interpretability · The AI Lecture Hall
- Random Forests: Decorrelated Trees and Robust Predictions · The AI Lecture Hall
- Bagging and the Bootstrap: Variance Reduction by Averaging · The AI Lecture Hall
- Boosting II: Gradient Boosting Machines · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/decision-trees-and-random-forests.html