An experimental study of diffusion-based general speech restoration with predictive-guided conditioning

Citations

WEB OF SCIENCE

1
Citations

SCOPUS

1

초록

This study presents a hybrid speech restoration framework that integrates predictive-guided conditioning into a diffusion-based generative model to address complex distortions, including noise, reverberation, and bandwidth reduction. The proposed method employs the outputs of a predictive model to guide the diffusion process, enabling more accurate reconstruction under challenging acoustic conditions. Furthermore, during the final sampling stage, the outputs of the predictive and generative models are fused with a tunable ratio, balancing signal fidelity and perceptual naturalness. Experimental results demonstrate that the proposed approach significantly improves objective restoration metrics compared to conventional diffusion baselines. However, the perceptual quality varies with the fusion ratio, revealing a trade-off between objective gains and subjective preference. These findings highlight the potential of predictive-guided conditioning for robust speech restoration and provide insights into optimizing the balance between predictive and generative contributions.

키워드

Score-based diffusion modelPredictive-guided conditioningGeneral speech restorationSpeech enhancementAcoustic noiseArchitectural acousticsDiffusionRestorationSpeech communicationSpeech enhancement
제목
An experimental study of diffusion-based general speech restoration with predictive-guided conditioning
저자
Yang, Da-HeeChang, Joon-Hyuk
DOI
10.1016/j.csl.2026.101940
발행일
2026-07
유형
Article
저널명
Computer Speech and Language
99
페이지
1 ~ 11