Image Segmentation: Every Pixel Labelled
Semantic segmentation labels every pixel by class; instance segmentation separates each object. Plus the U-Net architecture that made it practical.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. Boxes are rough. Sometimes we need the exact outline of every object: for self-driving cars, medical scans and photo editing. That is segmentation.
Semantic segmentation. Semantic segmentation gives every pixel a class: sky, road, car, person or tree. Watch the mask sweep across the scene. Notice that both cars share the same colour, because semantic segmentation only cares about the class, not which car is which.
Instance segmentation. Instance segmentation goes further: each object gets its own mask. Now the two cars are different colours, so we can count them and track them separately. Mask R-CNN is a well-known model for this.
U-Net. Many segmentation networks use an encoder-decoder shape like U-Net. The encoder shrinks the image to understand what is in it. The decoder expands it back to full resolution to say exactly where. Skip connections carry fine details straight across, so boundaries stay sharp.
Applications. Segmentation outlines tumours in medical scans, finds drivable road for self-driving cars, powers portrait mode and background removal, and maps crops and buildings from satellite images.
Recap. To recap. Semantic segmentation labels pixels by class. Instance segmentation separates objects. Encoder-decoder networks like U-Net do the work, and we measure quality with IoU per class.