Glossary › audioldm
GLOSSARY
audioldm
appears in 1 paper titles
Definition
A latent diffusion model for text-to-audio. Instead of denoising a waveform it compresses mel-spectrograms into a VAE latent space, runs diffusion there, and decodes with a vocoder — the move that made latent diffusion practical for images. Conditioning uses a contrastive audio–text embedding space, so audio without captions can still contribute to training, since the audio embedding stands in for the missing text. It targets sound effects and ambience rather than music.