AI in Motion

Data Augmentation for Vision

Computer VisionBeginner1:246 chapters

One image becomes many training examples: flips, rotations, crops, lighting changes and noise teach models what really matters.

πŸ“„ Illustrated notes Β· every chapter as a picture Β· printable

Shortcuts: Space play/pause Β· ←/β†’ 5 s Β· N/P chapter Β· M voice Β· C subtitles Β· F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 Why is flipping handwritten digits risky?
Q2 Data augmentation mainly helps to…
Q3 What does mixup do?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. Labelled images are expensive. Data augmentation creates new training examples from the ones you already have, and it is one of the most effective tricks in computer vision.

Transformations. Take one image and transform it. Flip it horizontally. Rotate it slightly. Crop a random region. Make it brighter or darker. Add a little noise. The label stays the same: it is still a house. The model learns that these changes do not matter.

Why it works. Augmentation teaches invariance, the idea that a flipped or darker house is still a house. It makes memorising exact pixels harder, which reduces overfitting. And it simulates the variety of the real world.

Pitfalls. Choose augmentations that keep the label true. Horizontal flips suit most photos, but flipping text or digits can change their meaning. Colour changes are harmful if colour is the very thing you are classifying.

Advanced. Advanced methods include cutout, which erases random patches, mixup, which blends two images and their labels, and RandAugment, which picks strong augmentations automatically.

Recap. To recap. Transform images but keep their labels. It teaches invariance and reduces overfitting. Pick transforms that keep the label true, and apply them on the fly during training.