PaperLens
紙
Students
Professional
JA
EN
◐
Sign in with Google
Sign in
Read
Home
Close reading
New
Textbook
Go deeper
Learn
Lab
Landscape
Contributors
Glossary
You
Search
All-access
My Page
#llava
1 articles
01
2026-08-13
·
VLMs & Multimodal
·
★ MEMBER
·
PAPER
·
8 min read
How VLMs Came Together — Wiring a Vision Encoder into an LLM
A vision-language model is three parts: a vision encoder, a projection layer, and an LLM. From the bridge CLIP built, through the three schools of connector design, to resolution strategies and the implementation constraints that bite in practice.