A Momentum-Based Framework with Contrastive Data Generation for Robust Sound Source Localization

Citations

SCOPUS

0

초록

We propose MoCo-SSL, a momentum-based contrastive learning framework for multi-channel sound source localization (SSL) that enhances azimuth-aware representation learning. While prior SSL studies have used contrastive learning to handle varied acoustic conditions, we emphasize hard negatives-pairs with distinct azimuths recorded in the same room-for learning fine-grained spatial cues. A curriculum-based strategy gradually increases the proportion of such samples to raise task difficulty. The momentum contrast design employs a key encoder that maintains stable embeddings during curriculum transitions and receives audio with less noise and reverberation to produce clearer azimuth cues, thereby guiding the query encoder toward robust representations. Experiments show that MoCo-SSL consistently surpasses baselines, demonstrating the value of structured and noise-resilient representation learning in challenging SSL scenarios.

키워드

contrastive learningcurriculum learningexponentially moving averagemomentumsound source localizationAcoustic generatorsAcoustic noiseAcoustic noise measurementArchitectural acousticsAudio acousticsContrastive LearningCurricula
제목
A Momentum-Based Framework with Contrastive Data Generation for Robust Sound Source Localization
저자
Kim, Hyun-SooYang, Da-HeeChang, Joon-Hyuk
DOI
10.1109/ASRU65441.2025.11434773
발행일
2026-04
유형
Conference paper
저널명
ASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop
페이지
1 ~ 7