AI in Motion

Computer VisionIntermediate1:31 video6 chapters

Vision Transformers: Images as Patches — lecture notes

Cut an image into patches, treat them like words and let a Transformer attend between them. How ViTs work and when they beat CNNs.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Vision Transformers: Images as Patches

Transformers conquered language. In 2020, researchers asked: what if we treat an image like a sentence? The result was the Vision Transformer, or ViT.

0:112. Patches become tokens

Patches become tokens — Vision Transformers: Images as Patches

The image is cut into a grid of patches. Each patch is flattened and turned into an embedding, just like a word. A position number is added so the model knows where each patch came from, plus a special classification token. The Transformer then lets every patch attend to every other patch, and the classification token gives the answer.

0:353. CNNs vs ViTs

CNNs vs ViTs — Vision Transformers: Images as Patches

CNNs build in assumptions about local patterns, so they learn well from less data. Vision transformers have fewer built-in assumptions and see the whole image from the first layer. They shine when pre-trained on huge datasets, and they power many of today’s vision foundation models.

0:544. Attention

Attention — Vision Transformers: Images as Patches

Inside, it is the same attention we saw for words. Each patch builds its understanding by weighing every other patch, so distant parts of the image can inform each other directly.

1:085. Foundation models

Foundation models — Vision Transformers: Images as Patches

Vision transformers underpin CLIP, which connects images and text, self-supervised methods like DINO and MAE, promptable segmentation like Segment Anything, and multimodal assistants that can discuss images.

1:206. Recap

Recap — Vision Transformers: Images as Patches

To recap. Split into patches, embed them with positions, add a classification token, and let attention do the rest. Vision transformers shine with large-scale pre-training.

Key takeaways

  • ViTs split an image into patches and treat them as tokens.
  • Position embeddings and a [CLS] token are added before the Transformer.
  • Every patch can attend to every other patch from the first layer.
  • ViTs excel with large-scale pre-training and underpin many foundation models.

Check yourself

  1. In a ViT, what plays the role of words?
    Show answer

    Image patches — Each patch becomes one token.

  2. Why are position embeddings added?
    Show answer

    So the model knows where each patch came from — Attention alone ignores order and position.

  3. When do ViTs typically outperform CNNs?
    Show answer

    With very large datasets or pre-training — Fewer built-in assumptions need more data.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/vision-transformers.html