TSP-TTS: Text-based Style Predictor with Residual Vector Quantization for Expressive Text-to-Speech

Citations

WEB OF SCIENCE

2
Citations

SCOPUS

4

초록

Expressive text-to-speech (TTS) aims to synthesize better human-like speech by incorporating diverse speech styles or emotions. While most expressive TTS models rely on reference speech to condition the style of the generated speech, they often fail to generate speech of regular quality. To ensure consistent speech quality, we propose an expressive TTS conditioned on style representation extracted from the text itself. To implement this text-based style predictor, we design a style module incorporating residual vector quantization. Furthermore, the style representation is enhanced through style-to-text alignment and a mel decoder with style hierarchical layer normalization (SHLN). Our experimental findings demonstrate that our proposed model accurately estimates style representation, enabling the generation of high-quality speech without the need for reference speech.

키워드

expressive speech synthesisresidual vector quantizationText-to-speechConditionExpressive speech synthesisHuman likeResidual vector quantizationsSpeech emotionsSpeech modelsSpeech qualitySpeech styleText alignmentsText to speech
제목
TSP-TTS: Text-based Style Predictor with Residual Vector Quantization for Expressive Text-to-Speech
저자
Seong, DonghyunLee, HoyoungChang, Joon-Hyuk
DOI
10.21437/Interspeech.2024-1734
발행일
2024-09
유형
Proceedings Paper
저널명
INTERSPEECH 2024
페이지
1780 ~ 1784