AI in Motion

Transfer Learning for Vision

Computer VisionBeginner1:226 chapters

Reuse a network trained on millions of images: freeze its layers, add a new head, and fine-tune with only a small dataset of your own.

📄 Illustrated notes · every chapter as a picture · printable

Shortcuts: Space play/pause · ←/→ 5 s · N/P chapter · M voice · C subtitles · F fullscreen

Quick quiz

3 questions to check your understanding.

Q1 In feature extraction, which part is trained?
Q2 Why keep early layers frozen when fine-tuning?
Q3 What learning rate is typical for fine-tuning pretrained layers?

Go deeper

University-level written lectures in The AI Lecture Hall:

Transcript

Introduction. Training a vision model from scratch needs millions of labelled images. Transfer learning lets you reuse a network that has already learned to see.

Frozen layers and a new head. Take a network pretrained on ImageNet, over a million images. Its layers already detect edges, textures, parts and objects. Freeze them, remove the original classifier, and add a new head for your own classes. Only the head is trained.

Fine-tuning. With a bit more data, unfreeze the last block as well and fine-tune it with a small learning rate. Early layers stay frozen, because edges and textures are useful for almost any task.

Choosing. With very little data, freeze everything and train only the head. With more data, or data very different from the original, fine-tune some of the top layers too, gently.

Tips. Match the preprocessing the network was trained with. Use a small learning rate for pretrained layers. Augment your data. And consider modern foundation models, which make very strong starting points.

Recap. To recap. Start from a pretrained network. Freeze early layers and add a new head. Fine-tune the top layers when data allows. And get strong results even with small datasets.