W2V2-Light: A Lightweight Version of Wav2vec 2.0 for Automatic Speech Recognition

Citations

WEB OF SCIENCE

4
Citations

SCOPUS

6

초록

Wav2vec 2.0 (W2V2) has shown remarkable speech recognition performance by pre-training only with unlabeled data and fine-tuning with a small amount of labeled data. However, the practical application of W2V2 is hindered by hardware memory limitations, as it contains 317 million parameters. To address this issue, we propose W2V2-Light, a lightweight version of W2V2. We introduce two simple sharing methods to reduce the memory consumption as well as the computational costs of W2V2. Compared to W2V2, our model has 91% lesser parameters and a speedup of 1.31 times with minor degradation in downstream task performance. Moreover, by quantifying the stability of representations, we provide an empirical insight into why our model is capable of maintaining competitive performance despite the significant reduction in memory.

키워드

attention alignmentAutomatic speech recognitionparameter sharingrepresentation learningsemi-supervised learningSpeech recognitionSupervised learningSpeech communicationAttention alignmentAutomatic speech recognitionFine tuningLabeled dataParameter sharingPre-trainingRepresentation learningSemi-supervised learningSpeech recognition performanceUnlabeled data
제목
W2V2-Light: A Lightweight Version of Wav2vec 2.0 for Automatic Speech Recognition
저자
Kim, Dong-HyunLee, Jae-HongMo, Ji-HwanChang, Joon-Hyuk
DOI
10.21437/Interspeech.2022-10339
발행일
2022-09
유형
Proceedings Paper
저널명
INTERSPEECH 2022
2022-September
페이지
3038 ~ 3042