Tokenized Generative Speech Enhancement With Language Model and Flow Matching

Citations

WEB OF SCIENCE

2
Citations

SCOPUS

2

초록

We propose a novel generative speech enhancement (SE) framework that integrates a language model (LM) and a flow-matching model. To utilize an LM with discrete tokens, we introduce dMel, which discretizes Mel spectrograms into a predefined set of quantized values on a linear-scale without requiring additional neural networks. dMel preserves both semantic and acoustic characteristics, providing a compact and effective token-based alternative to Mel spectrograms. We design the first encoder-decoder LM for SE, which learns to map noisy dMel to enhanced ones. Subsequently, flow-matching de-quantizes enhanced dMel into continuous representation and refines it by learning the optimal transport-based probability path, improving perceptual quality. This unified approach enables structured reconstruction while effectively suppressing noise. Experimental results demonstrate the effectiveness of our method in enhancing speech quality, establishing a new paradigm for generative SE without reliance on neural codec-based representations.

키워드

SpectrogramNoise measurementSpeech enhancementTokenizationDecodingTrainingNoiseIndexesComputational modelingAcousticstokenizationlanguage modelflow-matchingComputational linguisticsNeural networksOptimizationSemanticsSpeech codingSpeech communication
제목
Tokenized Generative Speech Enhancement With Language Model and Flow Matching
저자
Yang, Da-HeeLee, JaeukChang, Joon-Hyuk
DOI
10.1109/LSP.2025.3589128
발행일
2025-07
유형
Article
저널명
IEEE Signal Processing Letters
32
페이지
2828 ~ 2832