AI in Motion

Image Classification End to End

Computer VisionBeginner1:206 chapters

From a labelled dataset to a trained classifier: the full pipeline, softmax probabilities and how to judge the results.

πŸ“„ Illustrated notes Β· every chapter as a picture Β· printable

Shortcuts: Space play/pause Β· ←/β†’ 5 s Β· N/P chapter Β· M voice Β· C subtitles Β· F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 What do softmax outputs add up to?
Q2 A model recognises cows only when there is grass in the picture. This is…
Q3 What is usually the best way to start a new image classifier?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. Image classification answers one question: what is in this picture? Let us walk through the whole pipeline, from data to prediction.

Picture to prediction. An image goes in. A convolutional network extracts features layer by layer. The final layer gives a score for each class, and softmax turns those scores into probabilities: here 86 percent house, 8 percent barn, 4 percent castle and 2 percent tent.

The pipeline. Collect and label images for every class. Split off validation and test sets. Augment the training images for variety. Train, usually by fine-tuning a pretrained network. Then evaluate, class by class, and study the mistakes.

Softmax. Softmax exponentiates each class score so it is positive, then divides by the total so everything sums to one. The largest score gets the largest probability.

Pitfalls. Watch for pitfalls. Imbalanced classes. Shortcut learning, where the model looks at the background instead of the object. Near duplicate images leaking into the test set. And overconfidence: a high probability is not a guarantee.

Recap. To recap. Image, features, scores, softmax. Collect, split, augment, train and evaluate. Start from a pretrained model, and always inspect the mistakes.