Abstract
Cultural bias remains a major challenge in multilingual large language models (MLLMs). While these models have revolutionized natural language processing by enabling cross-lingual understanding and generation, they often perpetuate cultural biases from their training data, particularly for low-resource languages. This study addresses this gap by focusing on MLLMs for Afaan Oromo and Amharic, two official Ethiopian languages, alongside English. In addition to their low-resource status, historical marginalization, skewed representation in global web corpora, and the underrepresentation of culturally specific linguistic expressions make them prone to exhibiting cultural bias. Unfortunately, research on these two languages has been neglected, and no relevant benchmark has been constructed. To evaluate cultural bias in these two languages, we construct benchmarks by translating the StereoSet dataset into both languages and have native speakers manually verify the translations to preserve cultural nuances and bias contexts. Our work systematically examines cultural bias in MLLMs, demonstrating that it is a universal issue across all models tested. We then propose systematic approaches to mitigate this bias, specifically by combining proactive data filtering with counterfactual data augmentation for balanced training data, and self-debiasing for realtime inference-level correction. This integrated approach achieves consistent 7–12% reductions in cultural bias across all tested models. Our work shows that balanced multilingual training combined with novel targeted debiasing techniques can significantly reduce cultural bias while maintaining strong language modeling performance. Our findings establish that proactive data curation and tailored mitigation strategies are essential for building more equitable and culturally aware MLLMs for low-resource languages.
| Original language | English |
|---|---|
| Article number | 20250307 |
| Journal | Data Intelligence |
| Volume | 8 |
| Issue number | 2 |
| DOIs | |
| Publication status | Published - 1 Jun 2026 |
Fingerprint
Dive into the research topics of 'Mitigating Cultural Bias for Low-Resource Languages in Multilingual Large Language Models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver