Multimodal Models and CLIP
Put images and text in one shared space by pulling matching pairs together — the idea behind CLIP, image search and vision-language assistants.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
3 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. The newest AI assistants can look at a photo and talk about it. The key idea is to put images and text in one shared space. CLIP, from 2021, showed how.
Contrastive learning. CLIP has an image encoder and a text encoder. For a batch of images and their captions, it computes the similarity of every image with every caption. Training pulls the matching pairs, on the diagonal, together, and pushes mismatched pairs apart. After training, a picture of a cat sits closest to the words a sleepy cat.
What it enables. A shared space enables a lot. Zero-shot classification, by comparing an image with a caption for each class. Image search with plain-language queries. Guiding image generators with text. And assistants that describe and reason about pictures.
Vision-language models. Many vision-language assistants connect a vision encoder to a language model. The image becomes a set of vectors, a projection layer maps them into the language model’s token space, and the model reads them alongside your question.
Recap. To recap. CLIP puts images and text in one space using contrastive training. That enables zero-shot classification and search, and connecting a vision encoder to a language model gives multimodal assistants.