Object detection explained
Understand boxes, IoU, NMS and mAP well enough to train, evaluate and debug a detector
A taste of a lesson
My detector finds individual shoppers well but misses people in crowded queues. Why?
A likely culprit is non maximum suppression. People in a queue overlap heavily, so the box for a person partly hidden behind another can exceed the NMS IoU threshold and be removed as a 'duplicate'. Check by looking at raw predictions before NMS: if the missing people appear there, raise the NMS threshold a little or try a crowd aware variant, and see how false duplicates change. Also check your labels: crowded scenes are often under labelled. Quick exercise: two boxes of area 100 overlap by 60. What is their IoU, and would an NMS threshold of 0.5 remove one?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Compute IoU and explain how predictions are matched to ground truth
- Describe one stage, two stage and set based detectors and their trade offs
- Explain non maximum suppression and when it removes real objects
- Read precision recall curves, AP and mAP under a stated protocol
- Improve detection of small objects and diagnose label quality problems
Lesson plan
- 1 Boxes and intersection over union Represent boxes correctly and compute IoU by hand. Start
- 2 How detectors produce boxes Compare one stage and two stage detectors and anchor based versus anchor free heads. Start
- 3 Non maximum suppression Remove duplicate boxes and understand where NMS goes wrong. Start
- 4 Set prediction and open vocabulary detection Understand detection transformers and text prompted detectors at a conceptual level. Start
- 5 Evaluating detectors properly Read precision recall curves, AP and mAP and compare results fairly. Start
- 6 Small objects and messy labels Diagnose common detection failures and choose targeted fixes. Start
Try asking
About this tutor
For learners who know image classification and want to understand how models find and box many objects in one image. You will compute intersection over union by hand, follow how detectors propose and score boxes, and see why non maximum suppression is needed and when it fails. The course compares one stage and two stage designs, anchor based and anchor free heads, and detection transformers that predict a set of objects directly, all at the level of ideas rather than code. Then you learn to read precision recall curves and mean average precision properly, handle small objects and messy labels, and judge open vocabulary detectors that find objects from text prompts.
Reviews
4.7
3 ratingsSample
- Ravi K.Sample
Tiling large drone images for small objects was exactly what I needed. Also found missing labels in my set thanks to the last lesson.
- Joao F.Sample
Looking at raw predictions before NMS showed my crowd problem immediately. The IoU exercises made the threshold settings make sense.
- Lena H.Sample
The mAP lesson cleared up why numbers from two papers could not be compared. Detection transformers went quickly, but the core idea came through.
About the teacher
Computer vision taught through real images, real failure cases and careful evaluation
9 tutors 406 lessons taught Sample
I teach computer vision: classification, detection, segmentation, document understanding, video and the newer models that combine images with language. Most of my work has been building vision systems that had to hold up outside the lab, under odd lighting, unusual cameras and labels that were not quite consistent. So my lessons spend as much time on data and evaluation as...
See Noor's profile and tutorsMore like this
Other tutors on the same or nearby topics.