Diagnosis-aware multitask fine-tuning of Whisper for dysarthric speech recognition

  • Chung, Yoona
  • Hong, Jeongmin
  • Lee, Jaehyuk
  • Kim, Eunchan
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

AbstractIndividuals with dysarthria exhibit irregular speech patterns that vary by disease, significantly reducing the accuracy of conventional speech-recognition systems. Previous studies have typically focused on a single disease group or used aggregated data without accounting for inter-disease variation, thereby limiting disease-specific insights. In this study, fluency metrics were extracted from a Korean dysarthric speech corpus across three disease groups (stroke, cerebral palsy, and peripheral neuropathy) and the diseases were classified based on these features. The performance of the disease-specific speech-recognition models was evaluated using the weighted character error rate (Weighted-CER). Results showed that classification based on fluency metrics achieved 99% accuracy. The disease-specific models improved the CER by up to 18.34 and 1.05 percentage points compared with the Whisper–Small model and a model trained on the entire dataset, respectively. In terms of Weighted-CER, the error rate decreased by up to 15.27 and 1.49 percentage points, respectively. These findings indicate that disease-specific models can meaningfully enhance speech recognition and underscore the importance of developing speech-recognition systems that can adapt to individual speech characteristics in patients with dysarthria

키워드

DysarthriaFluency metricsVoice qualityPathology fine-tuningASRAUTOMATIC SPEECHFUNCTION APPROXIMATIONPARAMETERSDISORDERSLANGUAGECHILDRENMODEL
제목
Diagnosis-aware multitask fine-tuning of Whisper for dysarthric speech recognition
저자
Chung, YoonaHong, JeongminLee, JaehyukKim, Eunchan
DOI
10.1016/j.specom.2026.103393
발행일
2026-05
유형
Article
저널명
Speech Communication
180
페이지
1 ~ 17