Multimodal Models and CLIP — lecture notes
Put images and text in one shared space by pulling matching pairs together — the idea behind CLIP, image search and vision-language assistants.
0:001. Introduction

The newest AI assistants can look at a photo and talk about it. The key idea is to put images and text in one shared space. CLIP, from 2021, showed how.
0:132. Contrastive learning

CLIP has an image encoder and a text encoder. For a batch of images and their captions, it computes the similarity of every image with every caption. Training pulls the matching pairs, on the diagonal, together, and pushes mismatched pairs apart. After training, a picture of a cat sits closest to the words a sleepy cat.
0:373. What it enables

A shared space enables a lot. Zero-shot classification, by comparing an image with a caption for each class. Image search with plain-language queries. Guiding image generators with text. And assistants that describe and reason about pictures.
0:534. Vision-language models

Many vision-language assistants connect a vision encoder to a language model. The image becomes a set of vectors, a projection layer maps them into the language model’s token space, and the model reads them alongside your question.
1:095. Recap

To recap. CLIP puts images and text in one space using contrastive training. That enables zero-shot classification and search, and connecting a vision encoder to a language model gives multimodal assistants.
Key takeaways
- CLIP trains an image encoder and a text encoder into one shared space.
- Contrastive learning pulls matching image–caption pairs together and pushes others apart.
- Shared spaces enable zero-shot classification and natural-language image search.
- Multimodal assistants connect a vision encoder to a language model.
Check yourself
- In CLIP’s similarity matrix, which cells should be highest after training?
Show answer
The diagonal — matching image–caption pairs — Matching pairs are pulled together.
- How does zero-shot classification with CLIP work?
Show answer
Compare the image with a text prompt for each class — The closest class description wins.
- What connects image features to a language model in many assistants?
Show answer
A projection layer into the token space — It maps vision features into the LLM’s input space.
Go deeper
- CLIP: Connecting Images and Language · The AI Lecture Hall
- Multimodal Models: Vision–Language and Beyond · The AI Lecture Hall
- Vision Transformers (ViT): Images as Sequences of Patches · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/multimodal-models-and-clip.html