Glossary › multimodal
GLOSSARY
multimodal
appears in 5 paper titles
Definition
Working with more than one type of signal — text, images, audio, video — inside a single model, typically by projecting each into a shared representation space. The point is not merely accepting several inputs but learning correspondences across them, so the model can answer questions about a picture or ground a caption in a video frame. Vision-language models are the most common instance today.
Explainers using this term
- Paper Walkthrough: Puffin-World — A World Model That Remembers Which Way Is UpPuffin-World: Scaling a Unified Multimodal Model with Native 3D World States
- Paper Explained: Video-DeepResearch — Agents That Watch a Video, Then Chase Down Every LeadVideo-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
- Paper Deep-Dive: The 'Physics' of Multimodal Pretraining — Which Way Does Knowledge Actually Flow?Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes