Glossary › multi-speaker
GLOSSARY
multi-speaker
appears in 1 paper titles
Definition
Covers more than one voice in a single system. In synthesis it means one model that can render any of many speakers, achieved by factoring speaker identity into an embedding kept separate from linguistic content — the same mechanism behind zero-shot voice cloning from a short reference clip. On the recognition side the phrase points instead at separating and attributing overlapping speech. The synthesis sense carries obvious consent and misuse concerns, which is why watermarking work travels alongside it.
Explainers using this term
- Paper Walkthrough: SwanTale — Designing Voices from Words Alone, with Speech and Sound in One WaveformSwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks