Cross-Modal Learning for CTC-Based ASR: Leveraging CTC-Bertscore and Sequence-Level Training

Citations

SCOPUS

1

초록

Due to the nature of neural networks that easily overfit the training set, neural network-based speech recognition models are vulnerable to prior shifts in data distribution or unseen words. Therefore, studies have been conducted to over-come this problem by using language models trained with a relatively easy-to-obtain unpaired corpus. In this paper, we present a new training method that uses BERT to improve the performance of a connectionist temporal classification (CTC)-based ASR model. The proposed method follows a cross-modal learning scenario and induces the CTC model to better embed contextual information by utilizing an auxiliary objective function operating at the sequence level. We applied the proposed method to fine-tune the pre-trained wav2vec 2.0 model with CTC loss and confirmed that the proposed method improves the generalization performance of the ASR model.

키워드

BERTConnectionist temporal classificationCross-modal learningSpeech recognitionBERTConnectionist temporal classificationCross-modalCross-modal learningData distributionNetwork-basedNeural-networksRecognition modelsTemporal classificationTraining sets
제목
Cross-Modal Learning for CTC-Based ASR: Leveraging CTC-Bertscore and Sequence-Level Training
저자
이문학이상언Choi, Ji-EunChang, Joon-Hyuk
DOI
10.1109/ASRU57964.2023.10389656
발행일
2023-12
유형
Conference paper
저널명
2023 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2023
페이지
1 ~ 8