Object Detection and Segmentation: A Deep Dive — lecture notes
Finding and outlining objects: sliding windows, R-CNN to Faster R-CNN, YOLO and one-stage detectors, anchors, IoU, non-maximum suppression, mAP, focal loss, DETR, semantic, instance and panoptic segmentation, U-Net and Segment Anything.
0:001. Introduction

Classifying a whole image is useful, but most real applications need more: where exactly is each car, each person, each tumour? In this deep dive we cover object detection, which draws boxes around objects, and segmentation, which labels every pixel. These tasks power self driving cars, medical imaging and much more.
0:212. The family of vision tasks

Vision tasks form a family. Classification gives one label per image. Detection gives a box and a label for every object. Semantic segmentation labels every pixel with a class, instance segmentation separates individual objects, and keypoint estimation finds landmarks such as joints for human pose.
0:403. Sliding windows

The classic approach slides a window across the image and asks a classifier: is there a car here? The score rises when the window covers a car. But a real image needs thousands of windows at many positions, sizes and shapes, which is very slow.
1:004. Two-stage detectors

The R CNN family made detection accurate. R CNN, in 2014, ran a CNN on about two thousand proposed regions per image: accurate but slow. Fast R CNN shared one CNN pass for the whole image, and Faster R CNN, in 2015, learned to propose regions too, with a small region proposal network.
1:225. Faster R-CNN

Here is Faster R CNN step by step. A backbone CNN turns the image into feature maps. A region proposal network scores many candidate boxes as object or background. Each promising region’s features are cropped to a fixed size, and a final head predicts its class and refines the box.
1:436. Anchors

Many detectors use anchor boxes: a set of predefined boxes of different sizes and shapes placed at every location of the feature map, such as tall ones for people and wide ones for cars. The network predicts small offsets from the best matching anchor, which is much easier than predicting boxes from nothing.
2:067. YOLO: look once

Modern real time detectors like YOLO look only once. The image is divided into a grid, and every cell predicts boxes and confidence scores for objects centred in it, all in a single forward pass. That produces many overlapping candidate boxes, which are cleaned up in a final step.
2:268. Two-stage vs one-stage

Two stage detectors propose regions and then classify them, which is accurate but slower. One stage detectors, such as YOLO, SSD and RetinaNet, predict boxes densely in a single pass, running at dozens of frames per second. Modern one stage designs have largely closed the accuracy gap.
2:469. Intersection over union

How do we measure whether a predicted box is right? Intersection over union: the overlap area divided by the combined area. Zero means no overlap and one means a perfect match. As the prediction slides into place, IoU rises to point six, and a common rule counts point five or more as correct.
3:0910. IoU

Intersection over union divides the area of overlap by the total area covered by either box. It is used to decide whether a detection counts as correct, to match predictions with ground truth during training, and to remove duplicate boxes, as we will see next.
3:2811. Compute it

Let us compute one. Two ten by ten boxes overlap in a five by ten strip. The intersection is fifty. The union is one hundred plus one hundred minus fifty, which is one hundred and fifty. So the IoU is one third, below the usual point five threshold.
3:4812. Non-maximum suppression

Detectors produce many overlapping boxes for the same object. Non maximum suppression cleans them up: sort boxes by confidence, keep the best one, delete other boxes of the same class that overlap it too much, and repeat. Dozens of guesses become one clean box per object.
4:0813. NMS by hand

Let us run non maximum suppression by hand. Three car boxes have confidences of point nine, point eight and point six, and the two weaker boxes each overlap the strongest one with an IoU above one half. Only the point nine box survives; the other two are suppressed as duplicates.
4:2914. Objects of every size

Objects appear at every size, from a distant pedestrian to a nearby bus. Feature pyramid networks combine deep, low resolution features, which know what an object is, with shallow, high resolution features, which know precisely where it is, at several scales. Small and large objects are then both detected well.
4:5115. Measuring detectors

Detectors are compared by mean average precision. For each class, precision and recall are traced as the confidence threshold varies, and the area under that curve is the average precision. Averaging over classes gives mAP. The COCO benchmark also averages over IoU thresholds from point five to point nine five, rewarding precise boxes.
5:1316. Focal loss

One stage detectors face extreme imbalance: almost all of their thousands of candidate boxes are easy background. RetinaNet introduced focal loss, which multiplies the usual cross entropy by one minus p to the power gamma, shrinking the loss for easy examples so training focuses on the hard ones.
5:3317. Transformers for detection

Transformers entered detection with DETR in 2020. It treats detection as predicting a set: a transformer outputs a fixed number of boxes, each matched one to one with a real object during training. That removes the need for hand designed anchors and non maximum suppression.
5:5318. Augmenting boxes

Data augmentation works for detection too, with one catch: when an image is flipped, cropped or scaled, every box and mask must be transformed in exactly the same way. Detectors such as YOLO also use mosaic augmentation, stitching four training images into one, so objects appear in varied contexts and sizes.
6:1419. Open-vocabulary detection

Classic detectors only know the classes they were trained on. Open vocabulary detectors combine detection with vision language models like CLIP, so you can ask for objects in plain text, such as every red backpack, including categories that were never labelled in the training data.
6:3320. Semantic segmentation

Now to segmentation. Semantic segmentation gives every pixel a class: sky, road, car, person or tree. Watch the mask sweep across the scene. Both cars share the same colour, because semantic segmentation only cares about the class of each pixel, not which object it belongs to.
6:5321. Instance segmentation

Instance segmentation goes further: each object gets its own mask. Now the two cars have different colours, so we can count them and track them separately. Mask R CNN, which adds a mask branch to Faster R CNN, is a well known model for this.
7:1222. U-Net

Many segmentation networks use an encoder decoder shape like U Net. The encoder shrinks the image to understand what is in it. The decoder expands it back to full resolution to say exactly where. Skip connections carry fine details straight across, so boundaries stay sharp. U Net was designed for medical images.
7:3423. Segmentation metrics

Segmentation is measured per class with IoU, also called the Jaccard index, or the closely related Dice score, twice the overlap divided by the total size of both masks, averaged over classes. Dice is also used as a training loss, especially in medical imaging where target structures are small.
7:5524. Pause and think

Pause and think. In a scan, a tumour covers one percent of the pixels, and a model labels every pixel healthy. Its pixel accuracy is ninety nine percent, yet its Dice score for the tumour is zero: it found nothing. That is why per class overlap scores are used.
8:1625. Panoptic segmentation

Panoptic segmentation unifies the two. Every pixel gets a class, and every countable object, called a thing, such as a car or a person, also gets its own instance identity, while amorphous regions called stuff, such as sky and road, do not. It gives a complete description of a scene.
8:3726. Segment Anything

In 2023, Meta’s Segment Anything model brought foundation models to segmentation. Click a point or draw a box, and it returns a mask for that object, even for kinds of objects it never saw during training. It was trained on a dataset of more than one billion masks.
8:5827. Keypoints and pose

A related task finds keypoints. Pose estimation networks predict a heatmap for every joint, a glowing blob showing where the left elbow or right knee is likely to be. The peak of each heatmap becomes a keypoint, and connecting them gives a skeleton that can be tracked frame by frame.
9:1928. Tracking

In video, detection becomes tracking. Objects are detected in every frame, and detections are linked across frames into tracks, using predicted motion and appearance features. Tracking lets systems count people crossing a line, follow players in sports analytics, or predict where a pedestrian will step next.
9:3829. Detection in code

Using a modern detector takes a few lines. Load a small YOLO model pre trained on the COCO dataset, run it on an image with a confidence threshold, and loop over the boxes. Each one has a class, a confidence and its corner coordinates. Non maximum suppression happens automatically.
9:5930. Applications

These techniques are everywhere. Self driving systems detect vehicles and pedestrians and segment the drivable road. Hospitals segment tumours and organs. Farms count fruit and spot weeds. Satellites map buildings and floods, and factories detect defects on production lines.
10:1631. Practical advice

Some practical advice. Annotation is expensive: boxes are much cheaper to draw than pixel masks, so choose the simplest task that meets your need. Fine tune a pre trained model. Check small objects and rare classes separately, and test under new conditions, such as night, rain or a different camera.
10:3732. Recap

To recap. Detection gives a box and class for every object, and segmentation labels every pixel. Two stage detectors propose then classify, while one stage detectors like YOLO predict in a single pass. IoU scores boxes, NMS removes duplicates and mAP compares detectors. Segmentation comes in semantic, instance and panoptic forms.
Key takeaways
- Detection predicts a box and label per object; segmentation labels every pixel (semantic, instance or panoptic).
- Two-stage detectors (R-CNN → Faster R-CNN) propose regions then classify; one-stage detectors (YOLO, SSD, RetinaNet) predict densely in one pass.
- IoU = overlap / union; a detection usually counts as correct at IoU ≥ 0.5.
- Non-maximum suppression removes duplicate boxes; mAP summarises precision–recall across classes (and IoU thresholds on COCO).
- Focal loss handles background imbalance; DETR predicts boxes as a set without anchors or NMS.
- U-Net’s encoder–decoder with skip connections is a standard segmentation design; SAM segments new objects from prompts.
Check yourself
- What does IoU measure?
Show answer
Overlap between two boxes divided by their combined area — Intersection over union.
- Two 10 × 10 boxes overlapping in a 5 × 10 strip have an IoU of about…
Show answer
0.33 — 50 / 150.
- What does non-maximum suppression do?
Show answer
Removes duplicate overlapping boxes, keeping the most confident — One clean box per object.
- Instance segmentation differs from semantic segmentation because it…
Show answer
Gives each object its own separate mask — Car #1 and car #2 get different masks.
- Why is pixel accuracy misleading for tumour segmentation?
Show answer
The dominant background class inflates it even if the tumour is missed — Use per-class IoU or Dice.
Go deeper
- Object Detection I: R-CNN, Fast R-CNN and Faster R-CNN · The AI Lecture Hall
- Object Detection II: YOLO and Real-Time Detection · The AI Lecture Hall
- Object Detection III: SSD, RetinaNet and the Focal Loss · The AI Lecture Hall
- Semantic Segmentation: FCN, U-Net and DeepLab · The AI Lecture Hall
- Vision Foundation Models: Segment Anything and Promptable Vision · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/detection-and-segmentation-deep-dive.html