JA EN
Glossary › vision-language-action

GLOSSARY

vision-language-action

appears in 2 paper titles

Definition

A model class for robotics that takes images plus a language instruction and emits robot actions directly, often by discretizing actions into tokens and predicting them the way text is predicted. The usual recipe starts from a pre-trained vision-language model and adds an action head, inheriting general visual and linguistic grounding. What separates it from a VLM is the output: control signals executed in the physical world, where a mistake cannot be undone by resampling.

Explainers using this term