Resolution Consistency Training on Time-Frequency Domain for Semi-Supervised Sound Event Detection

Citations

WEB OF SCIENCE

1
Citations

SCOPUS

1

초록

The fact that unlabeled data can be used for supervised learning is of considerable relevance concerning polyphonic sound event detection (PSED) because of the high costs of frame-wise labeling. While semi-supervised learning (SSL) for image tasks has been extensively developed, SSL for PSED has not been substantially explored due to data augmentation limitations. In this paper, we propose a novel SSL strategy for PSED called resolution consistency training (ResCT), combining unsupervised terms with the mean teacher using different resolutions of a spectrogram for data augmentation. The proposed method regularizes the consistency between the model predictions for different resolutions by controlling the sampling rate and window size. Experimental results show that ResCT outperforms other SSL methods on various evaluation metrics: event-f1 score, intersection-f1 score, and PSDSs. Finally, we report on some ablation studies for the weak and strong augmentation policies.

키워드

data augmentationmulti-resolutional trainingsemi-supervised learningsound event detectionFrequency domain analysisSpeech communicationSupervised learningPersonnel trainingData augmentationDifferent resolutionsF1 scoresMulti-resolutional trainingPolyphonic soundsSemi-supervisedSemi-supervised learningSound event detectionTime frequency domainUnlabeled data
제목
Resolution Consistency Training on Time-Frequency Domain for Semi-Supervised Sound Event Detection
저자
최원국Chang, Joon-Hyuk
DOI
10.21437/Interspeech.2023-350
발행일
2023-08
유형
Proceedings Paper
저널명
INTERSPEECH 2023
2023-August
페이지
286 ~ 290