JA EN
Glossary › audioldm

GLOSSARY

audioldm

appears in 1 paper titles

Definition

A latent diffusion model for text-to-audio. Instead of denoising a waveform it compresses mel-spectrograms into a VAE latent space, runs diffusion there, and decodes with a vocoder — the move that made latent diffusion practical for images. Conditioning uses a contrastive audio–text embedding space, so audio without captions can still contribute to training, since the audio embedding stands in for the missing text. It targets sound effects and ambience rather than music.