상세 보기
DNN based multi-speaker speech synthesis with temporal auxiliary speaker ID embedding
- Lee, Junmo;
- Song, Kwangsub;
- Noh, Kyoungjin;
- Park, Tae-Jun;
- Chang, Joon-Hyuk
Citations
WEB OF SCIENCE
1Citations
SCOPUS
2초록
In this paper, multi speaker speech synthesis using speaker embedding is proposed. The proposed model is based on Tacotron network, but post-processing network of the model is modified with dilated convolution layers, which used in Wavenet architecture, to make it more adaptive to speech. The model can generate multi speaker voice with only one neural network model by giving auxiliary input data, speaker embedding, to the network. This model shows successful result for generating two speaker's voices without significant deterioration of speech quality.
키워드
Deep learning; Multi speaker speech synthesis; Sequence to sequence; Speech synthesis; Deep learning; Deep neural networks; Deterioration; Embeddings; Auxiliary inputs; Neural network model; Post processing; Sequence to sequence; Significant deteriorations; Speaker id; Speech quality; Speech synthesis
- 제목
- DNN based multi-speaker speech synthesis with temporal auxiliary speaker ID embedding
- 저자
- Lee, Junmo; Song, Kwangsub; Noh, Kyoungjin; Park, Tae-Jun; Chang, Joon-Hyuk
- 발행일
- 2019-05
- 유형
- Conference Paper
- 저널명
- ICEIC 2019 - International Conference on Electronics, Information, and Communication
- 페이지
- 1 ~ 4