상세 보기
Short-Utterance Embedding Enhancement Method Based on Time Series Forecasting Technique for Text-Independent Speaker Verification
- Choi, Jeong-Hwan;
- Yang, Joon-Young;
- Chang, Joon-Hyuk
WEB OF SCIENCE
3SCOPUS
3초록
Short-utterance embedding, which is a speaker embedding extracted from a short utterance, shows poor speaker verification performance due to insufficient speaker information. To address the problem, we propose a method to map the set of short-utterance embeddings to a set of long-utterance embeddings based on a neural network. Specifically, a speech utterance is cropped into multiple segments whose durations are gradually increasing, and the speaker embeddings are extracted from the sequence of cropped segments using a pre-trained speaker embedding extractor. Subsequently, the sequence of embeddings is divided into a group of short-utterances embeddings and that of long-utterance embeddings. In our method, a sequence-to-sequence model based forecasting technique is employed, where an encoder transforms the group of short-utterance embeddings to a fixed-dimensional vector, and then a decoder converts the vector into a group of long-utterance embeddings. Experimental results on the VoxCeleb and Speakers in the Wild datasets show that our method improves the text-independent speaker verification performance under short utterance condition.
키워드
- 제목
- Short-Utterance Embedding Enhancement Method Based on Time Series Forecasting Technique for Text-Independent Speaker Verification
- 저자
- Choi, Jeong-Hwan; Yang, Joon-Young; Chang, Joon-Hyuk
- 발행일
- 2022-02
- 유형
- Proceedings Paper
- 저널명
- 2021 IEEE AUTOMATIC SPEECH RECOGNITION AND UNDERSTANDING WORKSHOP (ASRU)
- 페이지
- 130 ~ 137