Data augmentation for vision
Choose image augmentations that reflect real variation, keep labels correct and improve generalisation
A taste of a lesson
I am training a model to read gauge dials in a factory. Should I use random horizontal flips?
Probably not. A mirrored dial shows the needle on the opposite side, so a reading of 30 could look like 70: the transform changes the label, which teaches the model something false. Instead, think about what really varies in the factory: viewing angle, glare, dirt on the glass, lighting, camera blur. Small rotations, perspective changes, brightness and contrast jitter and mild blur fit those. Then look at a grid of augmented samples to check every reading is still correct. Quick check: would a large 180 degree rotation be safe for your gauges? Why or why not?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Choose geometric and photometric augmentations that match expected real variation
- Spot transforms that change the true label and exclude them
- Keep bounding boxes and masks consistent with transformed images
- Tune augmentation strength and judge its effect on validation data
- Decide when test time augmentation or synthetic images are worth the cost
Lesson plan
- 1 Why augmentation works Understand augmentation as teaching the model which changes should not matter. Start
- 2 Geometric and photometric transforms Know the main transform families and what real variation each one models. Start
- 3 When transforms break the label Identify augmentations that silently create wrong training examples. Start
- 4 Boxes, masks and mixing methods Augment detection and segmentation data correctly and use mixup and cutmix sensibly. Start
- 5 Strength, policies and testing Tune how strong augmentation is and measure whether it helps. Start
- 6 Test time augmentation and synthetic data Weigh the cost and benefit of augmenting at inference and generating images. Start
Try asking
About this tutor
For beginners training vision models who want to use augmentation deliberately rather than copying a default list. You will learn the main families of augmentation, geometric changes such as flips, crops and rotations and photometric changes such as brightness, colour and blur, and the key question behind each one: does this transform keep the label true? The course covers cases where it does not, such as mirrored text or left and right in medical scans, how to keep bounding boxes and masks in sync, mixing methods such as mixup and cutmix, choosing augmentation strength, test time augmentation and where synthetic images fit. Each lesson ends with a short decision exercise on a real scenario.
Reviews
4.0
3 ratingsSample
- Oliver N.Sample
Useful basics, but I hoped for more on augmentation for satellite images. The tutor gave general principles rather than domain specific recipes.
- Wei L.Sample
I had horizontal flips on for a model that reads signs with arrows. Removing them fixed a weird class confusion I had been chasing for a week.
- Carmen D.Sample
The 'what really varies in deployment' question is now my first step. The mixup section was a bit brief but enough to get started.
About the teacher
Computer vision taught through real images, real failure cases and careful evaluation
9 tutors 406 lessons taught Sample
I teach computer vision: classification, detection, segmentation, document understanding, video and the newer models that combine images with language. Most of my work has been building vision systems that had to hold up outside the lab, under odd lighting, unusual cameras and labels that were not quite consistent. So my lessons spend as much time on data and evaluation as...
See Noor's profile and tutorsMore like this
Other tutors on the same or nearby topics.