Mixup-Augmented Teacher-Guided Knowledge Distillation for Noise-Robust Keyword Spotting

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Keyword spotting (KWS) is a key component of voice interfaces and is typically deployed on resource-constrained devices. To improve the generalization performance of small-footprint models, knowledge distillation is widely adopted to transfer knowledge from an over-parameterized teacher to a compact student model. In addition, mixup-based data augmentation has been widely applied in KWS to enhance robustness through diverse training inputs. However, despite their potential effectiveness, the systematic integration of these two techniques remains underexplored in small-footprint KWS. We propose MixTKD (Mixup-Augmented Teacher-Guided Knowledge Distillation), an efficient distillation framework that leverages supervision from a teacher model to effectively integrate mixup for KWS. Specifically, teacher predictions serve as explicit guidance to formulate the cross-entropy and distillation losses. For the cross-entropy term, conventional mixup relies on linearly interpolated one-hot labels, which may introduce a discrepancy with the teacher output distribution used in the distillation term. To address this issue, MixTKD introduces teacher-guided soft targets that are aligned with the teacher outputs, providing semantically coherent supervision. For the distillation term, in contrast to conventional decoupled knowledge distillation, we perform a teacher-guided decomposition of the Kullback–Leibler divergence into target and non-target components based on teacher predictions, improving distillation effectiveness under mixed inputs. To further stabilize distillation, we propose teacher-confidence-based curriculum learning and a runner-up masking strategy. Experimental results show that the proposed method consistently outperforms conventional distillation methods on two KWS datasets, without increasing model size or inference latency. Furthermore, the method demonstrates robustness on noisy datasets and remains effective under limited distillation data.

키워드

Keyword spottingknowledge distillationmixup augmentationnoise robustnessArtificial intelligenceCurriculaEngineering researchEntropyKnowledge managementPersonnel trainingSpeech communicationStudentsTeaching
제목
Mixup-Augmented Teacher-Guided Knowledge Distillation for Noise-Robust Keyword Spotting
저자
Seong, Ju-SeokChang, Joon-Hyuk
DOI
10.1109/TASLPRO.2026.3703217
발행일
2026-06
유형
Article
저널명
IEEE Transactions on Audio, Speech and Language Processing
34
페이지
3344 ~ 3357