Image Segmentation: Every Pixel Labelled — lecture notes
Semantic segmentation labels every pixel by class; instance segmentation separates each object. Plus the U-Net architecture that made it practical.
0:001. Introduction

Boxes are rough. Sometimes we need the exact outline of every object: for self-driving cars, medical scans and photo editing. That is segmentation.
0:102. Semantic segmentation

Semantic segmentation gives every pixel a class: sky, road, car, person or tree. Watch the mask sweep across the scene. Notice that both cars share the same colour, because semantic segmentation only cares about the class, not which car is which.
0:283. Instance segmentation

Instance segmentation goes further: each object gets its own mask. Now the two cars are different colours, so we can count them and track them separately. Mask R-CNN is a well-known model for this.
0:434. U-Net

Many segmentation networks use an encoder-decoder shape like U-Net. The encoder shrinks the image to understand what is in it. The decoder expands it back to full resolution to say exactly where. Skip connections carry fine details straight across, so boundaries stay sharp.
1:015. Applications

Segmentation outlines tumours in medical scans, finds drivable road for self-driving cars, powers portrait mode and background removal, and maps crops and buildings from satellite images.
1:136. Recap

To recap. Semantic segmentation labels pixels by class. Instance segmentation separates objects. Encoder-decoder networks like U-Net do the work, and we measure quality with IoU per class.
Key takeaways
- Semantic segmentation assigns a class to every pixel.
- Instance segmentation also separates individual objects of the same class.
- U-Net uses an encoder, a decoder and skip connections for sharp masks.
- Applications include medical imaging, driving and photo editing.
Check yourself
- In semantic segmentation, two cars are…
Show answer
Given the same “car” label — Semantic segmentation labels classes, not instances.
- What do U-Net’s skip connections carry?
Show answer
Fine spatial detail from encoder to decoder — They help the decoder place boundaries precisely.
- Which task lets you count individual people in a crowd from masks?
Show answer
Instance segmentation — Each person gets a separate mask.
Go deeper
- Semantic Segmentation: FCN, U-Net and DeepLab · The AI Lecture Hall
- Instance Segmentation: Mask R-CNN and Beyond · The AI Lecture Hall
- Vision Foundation Models: Segment Anything and Promptable Vision · The AI Lecture Hall
- AI in Medical Imaging: Opportunities, Pitfalls and Validation · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/image-segmentation.html