Object Detection: Boxes, IoU and NMS
Find every object and draw a box around it. From sliding windows to YOLO-style grids, IoU and non-maximum suppression.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. Classification says what is in an image. Object detection says what and where: it draws a labelled box around every object.
Sliding windows. The classic approach slides a window across the image and asks a classifier: is there a car here? The score rises when the window covers a car. But a real image needs thousands of windows at many sizes, which is very slow.
Modern detectors. Modern detectors like YOLO look once. The image is divided into a grid, and every cell predicts boxes and confidence scores for objects centred in it, all in a single pass. That produces many overlapping candidate boxes. Non-maximum suppression keeps the best box for each object and removes the duplicates.
IoU. How do we measure whether a predicted box is right? Intersection over union: the overlap area divided by the combined area. Zero means no overlap, one means a perfect match. As the prediction slides into place, IoU rises to 0.6, and a common rule counts 0.5 or more as a correct detection.
Detector families. Two-stage detectors, like Faster R-CNN, first propose regions and then classify them: accurate but slower. One-stage detectors, like YOLO and SSD, predict everything in one pass, fast enough for real-time video.
Recap. To recap. Detection gives a class and a box for every object. Modern detectors predict many boxes at once. IoU measures overlap, and non-maximum suppression removes duplicates.