JA EN
Glossary › mllms

GLOSSARY

mllms

appears in 1 paper titles

Definition

Multimodal large language models: LLMs extended to take images, audio, or video alongside text. The usual construction bolts a pretrained modality encoder onto a pretrained language model and trains a projection that maps its features into the language model's embedding space, so most of the reasoning still happens in the text backbone. You will also see LMM for the same class of system; the abbreviations are used interchangeably.

Explainers using this term