跳到主要导航 跳到搜索 跳到主要内容

Mitigating Cultural Bias for Low-Resource Languages in Multilingual Large Language Models

  • Beijing Institute of Technology
  • Mettu University
  • Xinjiang University

科研成果: 期刊稿件文章同行评审

摘要

Cultural bias remains a major challenge in multilingual large language models (MLLMs). While these models have revolutionized natural language processing by enabling cross-lingual understanding and generation, they often perpetuate cultural biases from their training data, particularly for low-resource languages. This study addresses this gap by focusing on MLLMs for Afaan Oromo and Amharic, two official Ethiopian languages, alongside English. In addition to their low-resource status, historical marginalization, skewed representation in global web corpora, and the underrepresentation of culturally specific linguistic expressions make them prone to exhibiting cultural bias. Unfortunately, research on these two languages has been neglected, and no relevant benchmark has been constructed. To evaluate cultural bias in these two languages, we construct benchmarks by translating the StereoSet dataset into both languages and have native speakers manually verify the translations to preserve cultural nuances and bias contexts. Our work systematically examines cultural bias in MLLMs, demonstrating that it is a universal issue across all models tested. We then propose systematic approaches to mitigate this bias, specifically by combining proactive data filtering with counterfactual data augmentation for balanced training data, and self-debiasing for realtime inference-level correction. This integrated approach achieves consistent 7–12% reductions in cultural bias across all tested models. Our work shows that balanced multilingual training combined with novel targeted debiasing techniques can significantly reduce cultural bias while maintaining strong language modeling performance. Our findings establish that proactive data curation and tailored mitigation strategies are essential for building more equitable and culturally aware MLLMs for low-resource languages.

源语言英语
文章编号20250307
期刊Data Intelligence
8
2
DOI
出版状态已出版 - 1 6月 2026

学术指纹

探究 'Mitigating Cultural Bias for Low-Resource Languages in Multilingual Large Language Models' 的科研主题。它们共同构成独一无二的学术指纹。

引用此