Guided conditioning with predictive network on score-based diffusion model for speech enhancement

Citations

WEB OF SCIENCE

5
Citations

SCOPUS

6

초록

Although diffusion-based speech enhancement (SE) models have emerged, they exhibit lower ability in noise removal than other predictive-based SE models. This reflects a trade-off between generative models, which are capable of producing more natural speech based on estimated target distribution, and predictive models, which are more effective in noise removal. To mitigate this trade-off, we propose a novel conditioning method for score-based diffusion models. The proposed approach involves guiding the diffusion model with a pretrained predictive model without joint training, thereby enabling enhanced speech to offer the proper direction to the diffusion model. The effectiveness of the proposed method is highlighted by outperforming the baseline method, with only half the number of sampling steps.

키워드

Speech enhancementscore-based diffusion modelsgenerative modelingpredictive modelingconditioningSpeech enhancement
제목
Guided conditioning with predictive network on score-based diffusion model for speech enhancement
저자
Kim, DailYang, Da-HeeKim, DonghyunChang, Joon-HyukYang, JaemoChoi, JeonghwanLee, MoaMoon, Han-gil
DOI
10.21437/Interspeech.2024-1545
발행일
2024-09
유형
Proceedings Paper
저널명
INTERSPEECH 2024
페이지
1190 ~ 1194