상세 보기
A Momentum-Based Framework with Contrastive Data Generation for Robust Sound Source Localization
- Kim, Hyun-Soo;
- Yang, Da-Hee;
- Chang, Joon-Hyuk
SCOPUS
0초록
We propose MoCo-SSL, a momentum-based contrastive learning framework for multi-channel sound source localization (SSL) that enhances azimuth-aware representation learning. While prior SSL studies have used contrastive learning to handle varied acoustic conditions, we emphasize hard negatives-pairs with distinct azimuths recorded in the same room-for learning fine-grained spatial cues. A curriculum-based strategy gradually increases the proportion of such samples to raise task difficulty. The momentum contrast design employs a key encoder that maintains stable embeddings during curriculum transitions and receives audio with less noise and reverberation to produce clearer azimuth cues, thereby guiding the query encoder toward robust representations. Experiments show that MoCo-SSL consistently surpasses baselines, demonstrating the value of structured and noise-resilient representation learning in challenging SSL scenarios.
키워드
- 제목
- A Momentum-Based Framework with Contrastive Data Generation for Robust Sound Source Localization
- 저자
- Kim, Hyun-Soo; Yang, Da-Hee; Chang, Joon-Hyuk
- 발행일
- 2026-04
- 유형
- Conference paper
- 저널명
- ASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop
- 페이지
- 1 ~ 7