Prior-free Guided TTS: An Improved and Efficient Diffusion-based Text-Guided Speech Synthesis

Citations

WEB OF SCIENCE

3
Citations

SCOPUS

4

초록

Recently, diffusion models have exhibited higher sample quality with guidance, such as classifier guidance and classifier-free guidance. However, these guidances have limitations: they require extra classifiers or joint training, and incur additional sampling cost. In this study, we propose prior-free guidance diffusion model and prior-free guided text-to-speech (PfGuided-TTS) that can generate a speech at a quality as high as other guidances without extra training resources and computational cost. PfGuided-TTS can generate higher human perceptual quality speech than the existing autoregressive (AR) and non-AR models, including diffusion-based TTS on LJSpeech. In addition, we provide a schematic describing why and how classifier- and prior-free guided scores produce high-fidelity samples.

키워드

diffusion modelguided scoretext-to-speechSpeech communicationSpeech synthesisAdditional samplingAuto-regressiveAutoregressive modellingComputational costsDiffusion modelGuided scorePerceptual qualityResource costsSample qualityText to speechDiffusion
제목
Prior-free Guided TTS: An Improved and Efficient Diffusion-based Text-Guided Speech Synthesis
저자
최원국Kim, So-Jeong김태호Chang, Joon-Hyuk
DOI
10.21437/Interspeech.2023-506
발행일
2023-08
유형
Proceedings Paper
저널명
INTERSPEECH 2023
2023-August
페이지
4289 ~ 4293