View article

[PDF] from isca-archive.org

Investigating RNN-based speech enhancement methods for noise-robust Text-to-Speech.

Authors

Cassia Valentini-Botinhao, Xin Wang, Shinji Takaki, Junichi Yamagishi

Publication date

2016/9/13

Conference

SSW

Pages

146-152

Description

The quality of text-to-speech (TTS) voices built from noisy speech is compromised. Enhancing the speech data before training has been shown to improve quality but voices built with clean speech are still preferred. In this paper we investigate two different approaches for speech enhancement to train TTS systems. In both approaches we train a recursive neural network (RNN) to map acoustic features extracted from noisy speech to features describing clean speech. The enhanced data is then used to train the TTS acoustic model. In one approach we use the features conventionally employed to train TTS acoustic models, ie Mel cepstral (MCEP) coefficients, aperiodicity values and fundamental frequency (F0). In the other approach, following conventional speech enhancement methods, we train an RNN using only the MCEP coefficients extracted from the magnitude spectrum. The enhanced MCEP features and the phase extracted from noisy speech are combined to reconstruct the waveform which is then used to extract acoustic features to train the TTS system. We show that the second approach results in larger MCEP distortion but smaller F0 errors. Subjective evaluation shows that synthetic voices trained with data enhanced with this method were rated higher and with similar to scores to voices trained with clean speech.

Total citations

Cited by 408

201820192020202120222023202412 22 51 63 102 91 65

Scholar articles

Investigating RNN-based speech enhancement methods for noise-robust Text-to-Speech.

C Valentini-Botinhao, X Wang, S Takaki, J Yamagishi - SSW, 2016