Glossary › vision
GLOSSARY
vision
appears in 4 paper titles
Definition
Computer vision — the field concerned with interpreting images and video, spanning classification, detection, segmentation, depth, and generation. Its centre of gravity has shifted toward vision-language models, where visual input is processed jointly with text rather than in isolation. The word names the machine's side of the problem, not human perception, though the two literatures borrow from each other.
Explainers using this term
- Paper Explained — LatentPress: Feeding Compressed Context Straight to a Frozen LLM, Neither as Text Nor as PixelsLatentPress: Context Compression Beyond Text and Vision
- Paper walkthrough: Qwen-Drive-1.0 — bolting 3D perception and planning onto a VLM without touching its architectureQwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
- Paper Explained: Beyond Data Scaling — Why the Backbone, Not the Trajectory Count, Decides Your VLA (VLAct)Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models