Vision Transformers: Images as Patches — lecture notes
Cut an image into patches, treat them like words and let a Transformer attend between them. How ViTs work and when they beat CNNs.
0:001. Introduction

Transformers conquered language. In 2020, researchers asked: what if we treat an image like a sentence? The result was the Vision Transformer, or ViT.
0:112. Patches become tokens

The image is cut into a grid of patches. Each patch is flattened and turned into an embedding, just like a word. A position number is added so the model knows where each patch came from, plus a special classification token. The Transformer then lets every patch attend to every other patch, and the classification token gives the answer.
0:353. CNNs vs ViTs

CNNs build in assumptions about local patterns, so they learn well from less data. Vision transformers have fewer built-in assumptions and see the whole image from the first layer. They shine when pre-trained on huge datasets, and they power many of today’s vision foundation models.
0:544. Attention

Inside, it is the same attention we saw for words. Each patch builds its understanding by weighing every other patch, so distant parts of the image can inform each other directly.
1:085. Foundation models

Vision transformers underpin CLIP, which connects images and text, self-supervised methods like DINO and MAE, promptable segmentation like Segment Anything, and multimodal assistants that can discuss images.
1:206. Recap

To recap. Split into patches, embed them with positions, add a classification token, and let attention do the rest. Vision transformers shine with large-scale pre-training.
Key takeaways
- ViTs split an image into patches and treat them as tokens.
- Position embeddings and a [CLS] token are added before the Transformer.
- Every patch can attend to every other patch from the first layer.
- ViTs excel with large-scale pre-training and underpin many foundation models.
Check yourself
- In a ViT, what plays the role of words?
Show answer
Image patches — Each patch becomes one token.
- Why are position embeddings added?
Show answer
So the model knows where each patch came from — Attention alone ignores order and position.
- When do ViTs typically outperform CNNs?
Show answer
With very large datasets or pre-training — Fewer built-in assumptions need more data.
Go deeper
- Vision Transformers (ViT): Images as Sequences of Patches · The AI Lecture Hall
- CLIP: Connecting Images and Language · The AI Lecture Hall
- Self-Supervised Vision: SimCLR, MoCo, DINO and Masked Autoencoders · The AI Lecture Hall
- Multimodal Models: Vision–Language and Beyond · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/vision-transformers.html