H4C-TTS: Leveraging Multi-Modal Historical Context for Conversational Text-to-Speech

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

1

초록

Conversational text-to-speech (TTS) aims to synthesize natural voices appropriate to a situation by considering the context of past conversations as well as the current text. However, analyzing and modeling the context of a conversation remains challenging. Most conversational TTS use the content of historical and recent conversations without distinguishing between them and often generate speech that does not fit the situation. Hence, we introduce a novel conversational TTS, H4C-TTS, that leverages multi-modal historical context to realize contextually appropriate natural speech synthesis. To facilitate conversational context modeling, we design a context encoder that incorporates historical and recent contexts and a multi-modal encoder that processes textual and acoustic inputs. Experimental results demonstrate that the proposed model significantly improves the naturalness and quality of speech in conversational contexts compared with existing conversational TTS.

키워드

conversational speech synthesismulti-modalText-to-speech'currentContext modelsConversational speechConversational speech synthesisMulti-modalNatural speechQuality of speechText to speech
제목
H4C-TTS: Leveraging Multi-Modal Historical Context for Conversational Text-to-Speech
저자
Seong, DonghyunChang, Joon-Hyuk
DOI
10.21437/Interspeech.2024-1480
발행일
2024-09
유형
Proceedings Paper
저널명
INTERSPEECH 2024
페이지
4933 ~ 4937