Enhancing Target-speaker Automatic Speech Recognition Using Multiple Speaker Embedding Extractors with Virtual Speaker Embedding

  • Seong, Ju-Seok
  • Choi, Jeong-Hwan
  • Jeoung, Ye-Rin
  • Kim, Ilseok
  • Chang, Joon-Hyuk
Citations

SCOPUS

1

초록

Target-speaker automatic speech recognition (TS-ASR) utilizes speaker embeddings to identify a target speaker in multi-talker environments. While high-performance speaker embedding extractors provide discriminative embeddings, their computational demands limit practical deployment. In this study, we present two novel methods that effectively utilize lightweight extractors to enhance TS-ASR performance. First, we propose a multiple embeddings modulation that effectively transfers comprehensive speaker information to the ASR module, thereby improving overall performance and robustness against embedding variations. Second, we present a virtual speaker embedding augmentation technique that synthesizes embeddings of unseen speakers, reducing dependence on specific extractors while enhancing independent contributions from each extractor. Experimental results on the Libri2Mix dataset demonstrate that our proposed methods achieve significant WER reductions compared to the baseline model.

키워드

lightweight speaker embedding extractorspeaker embeddingTarget-speaker automatic speech recognitionvirtual speaker embeddingSpeech communicationSpeech recognition
제목
Enhancing Target-speaker Automatic Speech Recognition Using Multiple Speaker Embedding Extractors with Virtual Speaker Embedding
저자
Seong, Ju-SeokChoi, Jeong-HwanJeoung, Ye-RinKim, IlseokChang, Joon-Hyuk
DOI
10.21437/Interspeech.2025-2486
발행일
2025-08
유형
Conference paper
저널명
Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
페이지
4918 ~ 4922