Glossary › mllms
GLOSSARY
mllms
appears in 1 paper titles
Definition
Multimodal large language models: LLMs extended to take images, audio, or video alongside text. The usual construction bolts a pretrained modality encoder onto a pretrained language model and trains a projection that maps its features into the language model's embedding space, so most of the reasoning still happens in the text backbone. You will also see LMM for the same class of system; the abbreviations are used interchangeably.
Explainers using this term
- Paper Walkthrough: Annotations as Rollouts — Dropping the Ground Truth Into the Group as a Ninth AnswerAnnotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs