Confusion Matrix, Precision, Recall and ROC
Accuracy alone can mislead. Build a confusion matrix, compute precision, recall and F1, then move the threshold and trace an ROC curve.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. A spam filter is 86 percent accurate. Is it good? That depends on how it is wrong. Let us open up the confusion matrix.
The confusion matrix. We test the filter on 100 emails. 42 spam emails are caught: true positives. 8 real emails are wrongly marked as spam: false positives. 6 spam emails slip through: false negatives. And 44 real emails are correctly left alone. From these four numbers we get accuracy 86 percent, precision 84 percent, recall 87.5 percent and F1 85.7 percent.
Precision and recall. Precision asks: when the model says spam, how often is it right? Recall asks: of all the spam, how much did it catch? F1 balances the two. And accuracy can be misleading when one class is rare.
Moving the threshold. The model gives each email a score. The threshold decides which scores count as spam. Move it left and recall rises, but precision falls. Move it right and precision rises, but recall falls. Tracing every threshold draws the ROC curve. The area under it, the AUC, is about 0.94 here.
Choosing. Which metric matters depends on the cost of mistakes. A spam filter should favour precision, so real email is not hidden. Disease screening should favour recall, so sick patients are not missed.
Recap. To recap. Four outcomes make the confusion matrix. Precision and recall answer different questions. The threshold trades one for the other. And the ROC curve and AUC summarise performance across all thresholds.