Far-end Reconstruction-Guided Separation-based Acoustic Echo Cancellation with Flow Matching

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Acoustic echo cancellation (AEC) suppresses echo in a microphone signal while preserving near-end speech. Most neural AEC systems directly predict the near-end speech. We propose a framework that separates the representation into near-end and echo streams and adds a far-end reconstructor using flow matching that leverages the estimated echo. The backbone operates in real-time and comprises three stages—fusion, split, and reconstruction: it fuses the microphone input with a far-end reference, splits features into near-end and echo streams, and reconstructs both using source-aware skip connections and cross-stream interaction. The reconstructor is trained via conditional flow matching to recover a time-aligned far-end reference and is used only during training, so inference introduces no additional parameters or latency. We further introduce far-end reconstruction (FR)–guided self-supervised pretraining, which needs no near-end or echo labels and provides label-free, indirect guidance to the echo stream via reconstruction. On a customized dataset with diverse nonlinearities and delays and on the AEC Challenge blind test set, our method outperforms strong baselines, indicating that two-stream modeling plus FR-based guidance improves robustness to echo variability while preserving near-end quality.

키워드

Acoustic echo cancellationconditional flow matchingjoint trainingnoise suppressionself-supervised pretrainingsource separationAcoustic noiseAudio acousticsAudio signal processingEcho suppressionMicrophonesSelf-supervised learningSpeech communicationSpurious signal noise
제목
Far-end Reconstruction-Guided Separation-based Acoustic Echo Cancellation with Flow Matching
저자
Kim, Gyeong-SuChang, Joon-Hyuk
DOI
10.1109/TASLPRO.2026.3698005
발행일
2026-05
유형
Article in press
저널명
Ieee Transactions on Audio Speech and Language Processing
34
페이지
3384 ~ 3397