JA EN

#mfcc

1 articles

01 ·Audio & Speech·★ MEMBER·PAPER·13 min read Representing Sound — Mel Spectrograms and Audio Tokens Why speech models never eat raw waveforms, and what they eat instead: the chain from short-time Fourier transform to the mel scale to the log to discrete tokens. Covers the window-length tradeoff, why MFCCs dropped the DCT, how acoustic and semantic tokens differ, and the config mismatches that silently wreck audio in production.