AI in Motion

Generative AIIntermediate1:22 video5 chapters

Multimodal Models and CLIP — lecture notes

Put images and text in one shared space by pulling matching pairs together — the idea behind CLIP, image search and vision-language assistants.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Multimodal Models and CLIP

The newest AI assistants can look at a photo and talk about it. The key idea is to put images and text in one shared space. CLIP, from 2021, showed how.

0:132. Contrastive learning

Contrastive learning — Multimodal Models and CLIP

CLIP has an image encoder and a text encoder. For a batch of images and their captions, it computes the similarity of every image with every caption. Training pulls the matching pairs, on the diagonal, together, and pushes mismatched pairs apart. After training, a picture of a cat sits closest to the words a sleepy cat.

0:373. What it enables

What it enables — Multimodal Models and CLIP

A shared space enables a lot. Zero-shot classification, by comparing an image with a caption for each class. Image search with plain-language queries. Guiding image generators with text. And assistants that describe and reason about pictures.

0:534. Vision-language models

Vision-language models — Multimodal Models and CLIP

Many vision-language assistants connect a vision encoder to a language model. The image becomes a set of vectors, a projection layer maps them into the language model’s token space, and the model reads them alongside your question.

1:095. Recap

Recap — Multimodal Models and CLIP

To recap. CLIP puts images and text in one space using contrastive training. That enables zero-shot classification and search, and connecting a vision encoder to a language model gives multimodal assistants.

Key takeaways

  • CLIP trains an image encoder and a text encoder into one shared space.
  • Contrastive learning pulls matching image–caption pairs together and pushes others apart.
  • Shared spaces enable zero-shot classification and natural-language image search.
  • Multimodal assistants connect a vision encoder to a language model.

Check yourself

  1. In CLIP’s similarity matrix, which cells should be highest after training?
    Show answer

    The diagonal — matching image–caption pairs — Matching pairs are pulled together.

  2. How does zero-shot classification with CLIP work?
    Show answer

    Compare the image with a text prompt for each class — The closest class description wins.

  3. What connects image features to a language model in many assistants?
    Show answer

    A projection layer into the token space — It maps vision features into the LLM’s input space.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/multimodal-models-and-clip.html