Transfer Learning for Vision — lecture notes
Reuse a network trained on millions of images: freeze its layers, add a new head, and fine-tune with only a small dataset of your own.
0:001. Introduction

Training a vision model from scratch needs millions of labelled images. Transfer learning lets you reuse a network that has already learned to see.
0:112. Frozen layers and a new head

Take a network pretrained on ImageNet, over a million images. Its layers already detect edges, textures, parts and objects. Freeze them, remove the original classifier, and add a new head for your own classes. Only the head is trained.
0:273. Fine-tuning

With a bit more data, unfreeze the last block as well and fine-tune it with a small learning rate. Early layers stay frozen, because edges and textures are useful for almost any task.
0:424. Choosing

With very little data, freeze everything and train only the head. With more data, or data very different from the original, fine-tune some of the top layers too, gently.
0:555. Tips

Match the preprocessing the network was trained with. Use a small learning rate for pretrained layers. Augment your data. And consider modern foundation models, which make very strong starting points.
1:086. Recap

To recap. Start from a pretrained network. Freeze early layers and add a new head. Fine-tune the top layers when data allows. And get strong results even with small datasets.
Key takeaways
- Transfer learning reuses features learned on large datasets like ImageNet.
- Feature extraction freezes the backbone and trains only a new head.
- Fine-tuning unfreezes top layers and uses a small learning rate.
- It enables strong models from small datasets.
Check yourself
- In feature extraction, which part is trained?
Show answer
Only the new head — The pretrained backbone stays frozen.
- Why keep early layers frozen when fine-tuning?
Show answer
Edges and textures are useful for almost any task — Low-level features transfer well across tasks.
- What learning rate is typical for fine-tuning pretrained layers?
Show answer
A small one — Small steps avoid destroying useful pretrained features.
Go deeper
- Transfer Learning and Fine-Tuning · The AI Lecture Hall
- Self-Supervised Vision: SimCLR, MoCo, DINO and Masked Autoencoders · The AI Lecture Hall
- CLIP: Connecting Images and Language · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/transfer-learning-for-vision.html