DNN based multi-speaker speech synthesis with temporal auxiliary speaker ID embedding

  • Lee, Junmo
  • Song, Kwangsub
  • Noh, Kyoungjin
  • Park, Tae-Jun
  • Chang, Joon-Hyuk
Citations

WEB OF SCIENCE

1
Citations

SCOPUS

2

초록

In this paper, multi speaker speech synthesis using speaker embedding is proposed. The proposed model is based on Tacotron network, but post-processing network of the model is modified with dilated convolution layers, which used in Wavenet architecture, to make it more adaptive to speech. The model can generate multi speaker voice with only one neural network model by giving auxiliary input data, speaker embedding, to the network. This model shows successful result for generating two speaker's voices without significant deterioration of speech quality.

키워드

Deep learningMulti speaker speech synthesisSequence to sequenceSpeech synthesisDeep learningDeep neural networksDeteriorationEmbeddingsAuxiliary inputsNeural network modelPost processingSequence to sequenceSignificant deteriorationsSpeaker idSpeech qualitySpeech synthesis
제목
DNN based multi-speaker speech synthesis with temporal auxiliary speaker ID embedding
저자
Lee, JunmoSong, KwangsubNoh, KyoungjinPark, Tae-JunChang, Joon-Hyuk
DOI
10.23919/ELINFOCOM.2019.8706390
발행일
2019-05
유형
Conference Paper
저널명
ICEIC 2019 - International Conference on Electronics, Information, and Communication
페이지
1 ~ 4