JA EN
Glossary › multimodal

GLOSSARY

multimodal

appears in 5 paper titles

Definition

Working with more than one type of signal — text, images, audio, video — inside a single model, typically by projecting each into a shared representation space. The point is not merely accepting several inputs but learning correspondences across them, so the model can answer questions about a picture or ground a caption in a video frame. Vision-language models are the most common instance today.

Explainers using this term