Adversarial and Sequential Training for Cross-lingual Prosody Transfer TTS

Citations

WEB OF SCIENCE

1
Citations

SCOPUS

3

초록

This study presents a method for improving the performance of the text-to-speech (TTS) model by using three global speech-style representations: language, speaker, and prosody. Synthesizing different languages and prosody in the speaker's voice regardless of their own language and prosody is possible. To construct the embedding of each representation conditioned in the TTS model such that it is independent of the other representations, we propose an adversarial training method for the general architecture of TTS models. Furthermore, we introduce a sequential training method that includes rehearsal-based continual learning to train complex and small amounts of data without forgetting previously learned information. The experimental results show that the proposed method can generate good-quality speech and yield high similarity for speakers and prosody, even for representations that the speaker in the dataset does not contain.

키워드

adversarial trainingcontinual learningcross-lingualprosodytext-to-speechAdversarial trainingContinual learningCross-lingualPerformanceProsodyRepresentation languagesSpeech modelsSpeech styleText to speechTraining methodsSpeech communication
제목
Adversarial and Sequential Training for Cross-lingual Prosody Transfer TTS
저자
Kim, Min-KyungChang, Joon-Hyuk
DOI
10.21437/Interspeech.2022-865
발행일
2022-09
유형
Proceedings Paper
저널명
INTERSPEECH 2022
2022-September
페이지
4556 ~ 4560