Glossary › multi-modal
GLOSSARY
multi-modal
appears in 1 paper titles
Definition
Handling more than one kind of input or output — images, text, audio, video, actions — within a single model. The hard part is not accepting extra inputs but aligning representations of fundamentally different signals so the model learns which parts correspond to which. It is worth separating from cross-modal, which describes going from one modality to another, as in text-to-image retrieval; multi-modal describes the capability of operating over several at once.
Explainers using this term
- Paper Walkthrough: WeMM-Embedding — Putting Text, Images and Video on One RulerWeMM-Embedding: WeChat Multi-Modal Embedding Technical Report