General-purpose Adversarial Training for Enhanced Automatic Speech Recognition Model Generalization

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

We present a new adversarial training method called General-purpose adversarial training (GPAT) that enhances the performance of automatic speech recognition models. In GPAT, we propose the followings: (1) a plausible adversarial examples converter (PAC); (2) a distribution matching regularization term (DM reg.). Compared to previous studies that directly compute gradients with respect to the input, PAC incorporates non-linearity to achieve performance improvement while eliminating the need for extra forward passes. Furthermore, unlike previous studies that use fixed norms, GPAT can generate similar yet diverse samples through DM reg. We demonstrate that the GPAT elevates the performance of various models on the LibriSpeech dataset. Specifically, by applying GPAT to the conformer model, we achieved 5.3% average relative improvements. With respect to the wav2vec 2.0 experiments, our method yielded a 2.0%/4.4% word error rate on the LibriSpeech test sets without a language model.

키워드

adversarial trainingdata augmentationspeech recognitionSpeech communicationAdversarial trainingAutomatic speech recognitionData augmentationDistribution matchingModel generalizationPerformanceRecognition modelsRegularization termsTraining methodsWord error rateSpeech recognition
제목
General-purpose Adversarial Training for Enhanced Automatic Speech Recognition Model Generalization
저자
Kim, DoheeShim, DaeyeolChang, Joon-Hyuk
DOI
10.21437/Interspeech.2023-2389
발행일
2023-08
유형
Proceedings Paper
저널명
INTERSPEECH 2023
2023-August
페이지
889 ~ 893