跳到主要导航 跳到搜索 跳到主要内容

Monolingual anchoring for low-resource cross-lingual semantic alignment: A case study on uyghur

  • Ruohao Yan
  • , Huaping Zhang
  • , Yuwen Niu
  • , Jihong Zhu
  • , Askar Hamdulla*
  • *此作品的通讯作者
  • Xinjiang University
  • Beijing Institute of Technology
  • Tsinghua University

科研成果: 期刊稿件文章同行评审

摘要

Low-resource languages remain challenging for cross-lingual semantic alignment because of limited parallel corpora. In addition, conventional symmetric alignment may distort the semantic space of a high-resource language through noisy low-resource updates. To address this issue, we propose Monolingual Anchoring for Cross-Lingual Semantic Alignment (MACA), focusing on Uyghur as a low-resource case study. MACA follows an asymmetric paradigm that treats the high-resource language as a fixed semantic anchor and transfers its semantic structure to the Uyghur side. The method consists of three components: (1) Anchored Embedding Initialization for newly introduced Uyghur subwords, (2) Cross-Lingual Neighborhood Anchoring for structural alignment between Uyghur and the anchor language, and (3) Monolingual Structure Anchoring for improving the internal semantic organization of Uyghur representations. Experiments centered on Uyghur-Chinese show that MACA outperforms LaBSE, the strongest off-the-shelf multilingual baseline in our comparison, by 7.67 points on cross-lingual STS. In an exploratory Uyghur-English zero-shot setting, MACA also surpasses LaBSE by 2.45 points without using Uyghur-English training data. These results provide evidence for the effectiveness of MACA in the evaluated Uyghur setting and suggest that monolingual anchoring may be further explored for related low-resource languages, such as Kazakh, Kyrgyz, and Uzbek.

源语言英语
期刊论文编号736
期刊Journal of King Saud University - Computer and Information Sciences
38
7
DOI
出版状态已出版 - 9月 2026

学术指纹

探究 'Monolingual anchoring for low-resource cross-lingual semantic alignment: A case study on uyghur' 的科研主题。它们共同构成独一无二的学术指纹。

引用此