Efficient Trainable Front-Ends for Neural Speech Enhancement

Jonah Casebeer,Umut Isik,Shrikant Venkataramani,Arvindh Krishnaswamy

Efficient Trainable Front-Ends for Neural Speech Enhancement

2020

Jonah Casebeer
Umut Isik
Shrikant Venkataramani
Arvindh Krishnaswamy

Many neural speech enhancement and source separation systems operate in the time-frequency domain. Such models often benefit from making their Short-Time Fourier Transform (STFT) front-ends trainable. In current literature, these are implemented as large Discrete Fourier Transform matrices; which are prohibitively inefficient for low-compute systems. We present an efficient, trainable front-end based on the butterfly mechanism to compute the Fast Fourier Transform, and show its accuracy and efficiency benefits for low-compute neural speech enhancement models. We also explore the effects of making the STFT window trainable.

Keywords:

Discrete Fourier transform
Matrix (mathematics)
Fourier transform
Pattern recognition
Computer science
Short-time Fourier transform
Fast Fourier transform
Artificial intelligence
Speech enhancement
Source separation

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations