Mitigating Linguistic Bias Between Malay and Indonesian Languages Using Masked Language Models

  • Bit, Ferdinand Lenchau
  • binti Zamri, Iman Khaleda
  • Wasi, Amzine Toushik
  • Rafi, Taki Hasan
  • Chae, Dong-Kyu
Citations

SCOPUS

0

초록

Language models (LMs) are essential for natural language processing (NLP) tasks, but they often exhibit biases due to inadequate or imbalanced training data, particularly in multilingual settings. These biases can lead to challenges in modeling linguistically similar low-resource languages, such as Malay and Indonesian, where mutual intelligibility complicates language differentiation. Addressing these biases is critical for enhancing the performance and fairness of NLP tools for underrepresented languages. Current LMs struggle with consistency in such scenarios, often leading to language mixing or poor prediction accuracy due to insufficient data capturing subtle linguistic differences. To tackle this, we curate a novel dataset of Malay sentences infused with Indonesian intrusions by simulating mixed-language sentences through filtering and refinement. We then fine-tune a RoBERTa model on this dataset. Empirically, this model exhibits significant improvements in word-level accuracy and language consistency compared to baseline models, indicating its ability to mitigate biases effectively. We believe that our work offers a pathway to address linguistic gaps and fosters the development of more accurate and equitable NLP tools for low-resource languages in multilingual environments.

키워드

Bias in NLPMalay and Indonesian language modelsNLP for underrepresented languagesComputational linguisticsData accuracyNatural language processing systems
제목
Mitigating Linguistic Bias Between Malay and Indonesian Languages Using Masked Language Models
저자
Bit, Ferdinand Lenchaubinti Zamri, Iman KhaledaWasi, Amzine ToushikRafi, Taki HasanChae, Dong-Kyu
DOI
10.1007/978-981-95-3827-0_23
발행일
2026-06
유형
Conference Paper
저널명
Lecture Notes in Computer Science
15986
페이지
328 ~ 338