JA EN
Glossary › text-to-speech

GLOSSARY

text-to-speech

appears in 1 paper titles

Definition

The spelled-out form of TTS: converting written text into spoken audio. A full pipeline normalizes text and derives phonemes and prosody, predicts an acoustic representation, then renders a waveform. What makes it hard is not the symbol mapping but everything underdetermined by the text — rhythm, emphasis, speaker identity, emotion — which is why naturalness and speaker similarity are evaluated alongside plain intelligibility.