Image segmentation: semantic, instance, panoptic
Label every pixel correctly: understand the three kinds of segmentation, their models and metrics
A taste of a lesson
My road crack segmentation model has 96 percent pixel accuracy, but it barely finds any cracks. How?
Cracks are thin, so they might cover only around 3 or 4 percent of the pixels. A model that predicts 'no crack' everywhere already scores about 96 percent pixel accuracy, so that number tells you almost nothing. Measure IoU or Dice for the crack class alone instead. To help the model learn the small class, try Dice loss or a weighted cross entropy, and check that thin cracks survive any downsampling in your labels and inputs. Quick exercise: if cracks are 4 percent of pixels and the model predicts none, what is the IoU for the crack class?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Distinguish semantic, instance and panoptic segmentation and choose one for a task
- Explain encoder decoder and mask head approaches at a conceptual level
- Compute IoU and Dice from pixel counts and read mean IoU correctly
- Handle pixel level class imbalance and judge boundary quality
- Plan annotation using promptable segmentation models with human review
Lesson plan
- 1 Three kinds of segmentation Define semantic, instance and panoptic segmentation using one street scene. Start
- 2 How models produce masks Understand encoder decoder networks and detector based mask heads. Start
- 3 IoU, Dice and mean IoU Compute segmentation metrics by hand and compare results fairly. Start
- 4 Imbalance and boundaries Deal with dominant background pixels and judge edge quality. Start
- 5 Promptable models and annotation Use click and box prompted segmentation to label data faster while keeping quality. Start
- 6 Choosing an approach for a project Plan a segmentation project from task definition to evaluation. Start
Try asking
About this tutor
For learners who know classification and detection and want to work at pixel level. You will learn exactly how semantic, instance and panoptic segmentation differ and which one a project needs, then how models produce masks: encoder decoder networks with skip connections, and detectors with a mask head. You will compute IoU and Dice for a small example, see why background pixels dominate the loss and the metric, and learn why boundary errors can matter more than the scores suggest. The final lessons cover promptable segmentation models that produce masks from clicks or boxes, and how to plan annotation, which is often the most expensive part of a segmentation project.
Reviews
4.3
3 ratingsSample
- Mark T.Sample
Clear metric lessons. The tip to never compare Dice with IoU saved me from a misleading slide in a team meeting.
- Sana Q.Sample
Promptable segmentation for drafting labels cut our annotation time a lot. The tutor was right that the masks still needed checking at the edges.
- Elif Y.Sample
The street scene example made the three types click instantly. I realised my project needed instance masks, not semantic ones, before labelling anything.
About the teacher
Computer vision taught through real images, real failure cases and careful evaluation
9 tutors 406 lessons taught Sample
I teach computer vision: classification, detection, segmentation, document understanding, video and the newer models that combine images with language. Most of my work has been building vision systems that had to hold up outside the lab, under odd lighting, unusual cameras and labels that were not quite consistent. So my lessons spend as much time on data and evaluation as...
See Noor's profile and tutorsMore like this
Other tutors on the same or nearby topics.