상세 보기
Tokenized Generative Speech Enhancement With Language Model and Flow Matching
- Yang, Da-Hee;
- Lee, Jaeuk;
- Chang, Joon-Hyuk
WEB OF SCIENCE
2SCOPUS
2초록
We propose a novel generative speech enhancement (SE) framework that integrates a language model (LM) and a flow-matching model. To utilize an LM with discrete tokens, we introduce dMel, which discretizes Mel spectrograms into a predefined set of quantized values on a linear-scale without requiring additional neural networks. dMel preserves both semantic and acoustic characteristics, providing a compact and effective token-based alternative to Mel spectrograms. We design the first encoder-decoder LM for SE, which learns to map noisy dMel to enhanced ones. Subsequently, flow-matching de-quantizes enhanced dMel into continuous representation and refines it by learning the optimal transport-based probability path, improving perceptual quality. This unified approach enables structured reconstruction while effectively suppressing noise. Experimental results demonstrate the effectiveness of our method in enhancing speech quality, establishing a new paradigm for generative SE without reliance on neural codec-based representations.
키워드
- 제목
- Tokenized Generative Speech Enhancement With Language Model and Flow Matching
- 저자
- Yang, Da-Hee; Lee, Jaeuk; Chang, Joon-Hyuk
- 발행일
- 2025-07
- 유형
- Article
- 권
- 32
- 페이지
- 2828 ~ 2832