상세 보기
Mitigating Linguistic Bias Between Malay and Indonesian Languages Using Masked Language Models
- Bit, Ferdinand Lenchau;
- binti Zamri, Iman Khaleda;
- Wasi, Amzine Toushik;
- Rafi, Taki Hasan;
- Chae, Dong-Kyu
SCOPUS
0초록
Language models (LMs) are essential for natural language processing (NLP) tasks, but they often exhibit biases due to inadequate or imbalanced training data, particularly in multilingual settings. These biases can lead to challenges in modeling linguistically similar low-resource languages, such as Malay and Indonesian, where mutual intelligibility complicates language differentiation. Addressing these biases is critical for enhancing the performance and fairness of NLP tools for underrepresented languages. Current LMs struggle with consistency in such scenarios, often leading to language mixing or poor prediction accuracy due to insufficient data capturing subtle linguistic differences. To tackle this, we curate a novel dataset of Malay sentences infused with Indonesian intrusions by simulating mixed-language sentences through filtering and refinement. We then fine-tune a RoBERTa model on this dataset. Empirically, this model exhibits significant improvements in word-level accuracy and language consistency compared to baseline models, indicating its ability to mitigate biases effectively. We believe that our work offers a pathway to address linguistic gaps and fosters the development of more accurate and equitable NLP tools for low-resource languages in multilingual environments.
키워드
- 제목
- Mitigating Linguistic Bias Between Malay and Indonesian Languages Using Masked Language Models
- 저자
- Bit, Ferdinand Lenchau; binti Zamri, Iman Khaleda; Wasi, Amzine Toushik; Rafi, Taki Hasan; Chae, Dong-Kyu
- 발행일
- 2026-06
- 유형
- Conference Paper
- 권
- 15986
- 페이지
- 328 ~ 338