Textbook › Part II The Lineage of AI Models
CHAPTER 24
CLIP — Words and Images on One Map
There is one step you cannot skip on the way to understanding VLMs: CLIP, from 2021.
Until then, image recognition meant sorting a picture into one of a fixed set of classes decided in advance — a thousand of them, say. Anything the model had not been trained on simply could not be recognized. CLIP broke that frame.
What it did
Collect a huge number of image–caption pairs from the internet — roughly 400 million of them. Set up two encoders, one that processes images and one that processes text, and have each produce a vector.
Comments
Sign in to comment