Whisper Multilingual Downstream Task Tuning Using Task Vectors

Citations

WEB OF SCIENCE

3
Citations

SCOPUS

2

초록

Recently, the size of automatic speech recognition (ASR) models has been increasing, similar to large language models (LLMs), and efficient tuning to enhance the performance of downstream tasks with limited resources remains a challenge. In this paper, we propose a simple and effective downstream task tuning method using task vectors. We utilize task vectors to orient the pre-trained Whisper model in the weight space, moving in that direction to achieve downstream task adaptation. We demonstrate that the model can be adjusted through arithmetic operations of the task vector, and this adjustment is reflected in the Whisper. Furthermore, we can efficiently construct a generalized model by summing vectors. We set the direction of the model weight space for each multilingual language as the task vector to evaluate its effectiveness. We confirm that the task vector serves as a simple and effective approach for tuning downstream tasks in ASR using the Common Voice multilingual dataset.

키워드

downstream tasks adaptationmultilingualspeech recognitionCharacter recognitionSpeech enhancementSpeech recognitionVector spaces
제목
Whisper Multilingual Downstream Task Tuning Using Task Vectors
저자
Kang, Ji-HunLee, Jae-HongLee, Mun-HakChang, Joon-Hyuk
DOI
10.21437/Interspeech.2024-513
발행일
2024-09
유형
Proceedings Paper
저널명
INTERSPEECH 2024
페이지
2385 ~ 2389