JA EN
Textbook › Part II The Lineage of AI Models
CHAPTER 24

CLIP — Words and Images on One Map

★ MEMBER1 min

There is one step you cannot skip on the way to understanding VLMs: CLIP, from 2021.

Until then, image recognition meant sorting a picture into one of a fixed set of classes decided in advance — a thousand of them, say. Anything the model had not been trained on simply could not be recognized. CLIP broke that frame.

What it did

Collect a huge number of image–caption pairs from the internet — roughly 400 million of them. Set up two encoders, one that processes images and one that processes text, and have each produce a vector.

§

Members-only from here

All 26 chapters and every lab, $4.99/mo. Cancel anytime.

Comments

Sign in to comment