Extending Self-Distilled Self-Supervised Learning For Semi-Supervised Speaker Verification

Citations

SCOPUS

2

초록

In this study, we extend self-distillation with no labels (DINO), a successful self-supervised learning framework, by combining it with supervised classification (SC) for semi-supervised speaker verification with limited labeled data. We introduce a transfer learning framework that pre-trains and fine-tunes the encoder using DINO and SC, respectively, and a multitask learning framework that shares the encoder while having separate projection layers for both methods. To achieve lower inter-speaker similarity, we propose a joint learning framework sharing both the encoder and projection layer for DINO and SC. We also propose an auxiliary contrastive loss between embeddings derived from labeled and unlabeled utterances and introduce a two-stage learning strategy to apply margin penalty effectively. Experimental results on the VoxCeleb corpus indicate that the joint learning framework outperforms the other frameworks and is closest to achieving the performance of fully supervised learning.

키워드

DINOjoint learningself-supervised learningsemi-supervised learningspeaker verificationDistillationLearning systemsSemi-supervised learningSignal encoding
제목
Extending Self-Distilled Self-Supervised Learning For Semi-Supervised Speaker Verification
저자
최정환경제현성주석정예린Chang, Joon-Hyuk
DOI
10.1109/ASRU57964.2023.10389802
발행일
2023-12
유형
Conference paper
저널명
2023 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2023
페이지
1 ~ 8