Intra-ensemble: A New Method for Combining Intermediate Outputs in Transformer-based Automatic Speech Recognition

Citations

WEB OF SCIENCE

1
Citations

SCOPUS

1

초록

Deep learning models employ various regularization techniques to prevent overfitting and enhance generalization. In particular, an auxiliary loss, as proposed for connectionist temporal classification (CTC) models, demonstrated the potential for intermediate prediction to be useful by enabling sub-models to recognize speech accurately. We propose a new method called Intra-ensemble, which combines these accurate intermediate outputs into a single output for both training and inference, considering the importance of the intermediate layer using learnable parameters. Our approach is applicable to CTC models, attention-based encoder-decoder models, and transducer structures and demonstrated performance improvements of 13.5%, 3.0%, and 4.1% respectively, in the LibriSpeech evaluation. Furthermore, through various analytical experiments, we found that the sub-models contributed significantly to performance improvement.

키워드

ensemblespeech recognitionDeep learningSpeech communicationAutomatic speech recognitionClassification modelsEnsembleGeneralisationLearning modelsOverfittingPerformanceRegularization techniqueSubmodelsTemporal classificationSpeech recognition
제목
Intra-ensemble: A New Method for Combining Intermediate Outputs in Transformer-based Automatic Speech Recognition
저자
Kim, DoHeeChoi, JieunChang, Joon-Hyuk
DOI
10.21437/Interspeech.2023-1255
발행일
2023-08
유형
Proceedings Paper
저널명
INTERSPEECH 2023
2023-August
페이지
2203 ~ 2207