TY - JOUR
T1 - Mitigating Cultural Bias for Low-Resource Languages in Multilingual Large Language Models
AU - Binegde, Geleta Negasa
AU - Zhang, Huaping
AU - Li, Qiuchi
AU - Zhang, Baohua
AU - Gao, Chunxiao
AU - Yan, Ruohao
N1 - Publisher Copyright:
© 2026 Science Press. All rights reserved.
PY - 2026/6/1
Y1 - 2026/6/1
N2 - Cultural bias remains a major challenge in multilingual large language models (MLLMs). While these models have revolutionized natural language processing by enabling cross-lingual understanding and generation, they often perpetuate cultural biases from their training data, particularly for low-resource languages. This study addresses this gap by focusing on MLLMs for Afaan Oromo and Amharic, two official Ethiopian languages, alongside English. In addition to their low-resource status, historical marginalization, skewed representation in global web corpora, and the underrepresentation of culturally specific linguistic expressions make them prone to exhibiting cultural bias. Unfortunately, research on these two languages has been neglected, and no relevant benchmark has been constructed. To evaluate cultural bias in these two languages, we construct benchmarks by translating the StereoSet dataset into both languages and have native speakers manually verify the translations to preserve cultural nuances and bias contexts. Our work systematically examines cultural bias in MLLMs, demonstrating that it is a universal issue across all models tested. We then propose systematic approaches to mitigate this bias, specifically by combining proactive data filtering with counterfactual data augmentation for balanced training data, and self-debiasing for realtime inference-level correction. This integrated approach achieves consistent 7–12% reductions in cultural bias across all tested models. Our work shows that balanced multilingual training combined with novel targeted debiasing techniques can significantly reduce cultural bias while maintaining strong language modeling performance. Our findings establish that proactive data curation and tailored mitigation strategies are essential for building more equitable and culturally aware MLLMs for low-resource languages.
AB - Cultural bias remains a major challenge in multilingual large language models (MLLMs). While these models have revolutionized natural language processing by enabling cross-lingual understanding and generation, they often perpetuate cultural biases from their training data, particularly for low-resource languages. This study addresses this gap by focusing on MLLMs for Afaan Oromo and Amharic, two official Ethiopian languages, alongside English. In addition to their low-resource status, historical marginalization, skewed representation in global web corpora, and the underrepresentation of culturally specific linguistic expressions make them prone to exhibiting cultural bias. Unfortunately, research on these two languages has been neglected, and no relevant benchmark has been constructed. To evaluate cultural bias in these two languages, we construct benchmarks by translating the StereoSet dataset into both languages and have native speakers manually verify the translations to preserve cultural nuances and bias contexts. Our work systematically examines cultural bias in MLLMs, demonstrating that it is a universal issue across all models tested. We then propose systematic approaches to mitigate this bias, specifically by combining proactive data filtering with counterfactual data augmentation for balanced training data, and self-debiasing for realtime inference-level correction. This integrated approach achieves consistent 7–12% reductions in cultural bias across all tested models. Our work shows that balanced multilingual training combined with novel targeted debiasing techniques can significantly reduce cultural bias while maintaining strong language modeling performance. Our findings establish that proactive data curation and tailored mitigation strategies are essential for building more equitable and culturally aware MLLMs for low-resource languages.
UR - https://www.scopus.com/pages/publications/105044240720
U2 - 10.3724/2096-7004.di.2025.0307
DO - 10.3724/2096-7004.di.2025.0307
M3 - Article
AN - SCOPUS:105044240720
SN - 2096-7004
VL - 8
JO - Data Intelligence
JF - Data Intelligence
IS - 2
M1 - 20250307
ER -