Decision Trees and Random Forests
A tree learns yes/no questions that carve the data into pure regions. A forest of trees votes for more robust predictions.
๐ Illustrated notes ยท every chapter as a picture ยท printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. A doctor diagnoses by asking questions. A decision tree learns its own series of yes or no questions from data.
Growing a tree. Before any split, predicting pink for everything is right 60 percent of the time. The first question, is feature one below 5.5, raises accuracy to 77 percent. A second split makes it 87 percent, and a third makes every region pure: 100 percent on this training data.
Choosing splits. At each step, the tree tries many possible questions and measures how pure the resulting groups would be, using Gini impurity or entropy. It keeps the best question and repeats on each branch.
Overfitting. A tree grown too deep memorises noise. It scores perfectly on training data and badly on new data. We control it by limiting depth, requiring enough examples per leaf, or pruning.
Random forests. A random forest trains many trees, each on a random sample of the data and features. Each tree votes. Here four trees say cat and one says dog, so the forest predicts cat. Averaging many different trees makes predictions far more stable.
Recap. To recap. Decision trees learn readable questions. Splits aim for purity. Deep trees overfit. And random forests combine many trees into a strong, stable model.