Skip to main navigation Skip to search Skip to main content

Monolingual anchoring for low-resource cross-lingual semantic alignment: A case study on uyghur

  • Ruohao Yan
  • , Huaping Zhang
  • , Yuwen Niu
  • , Jihong Zhu
  • , Askar Hamdulla*
  • *Corresponding author for this work
  • Xinjiang University
  • Beijing Institute of Technology
  • Tsinghua University

Research output: Contribution to journalArticlepeer-review

Abstract

Low-resource languages remain challenging for cross-lingual semantic alignment because of limited parallel corpora. In addition, conventional symmetric alignment may distort the semantic space of a high-resource language through noisy low-resource updates. To address this issue, we propose Monolingual Anchoring for Cross-Lingual Semantic Alignment (MACA), focusing on Uyghur as a low-resource case study. MACA follows an asymmetric paradigm that treats the high-resource language as a fixed semantic anchor and transfers its semantic structure to the Uyghur side. The method consists of three components: (1) Anchored Embedding Initialization for newly introduced Uyghur subwords, (2) Cross-Lingual Neighborhood Anchoring for structural alignment between Uyghur and the anchor language, and (3) Monolingual Structure Anchoring for improving the internal semantic organization of Uyghur representations. Experiments centered on Uyghur-Chinese show that MACA outperforms LaBSE, the strongest off-the-shelf multilingual baseline in our comparison, by 7.67 points on cross-lingual STS. In an exploratory Uyghur-English zero-shot setting, MACA also surpasses LaBSE by 2.45 points without using Uyghur-English training data. These results provide evidence for the effectiveness of MACA in the evaluated Uyghur setting and suggest that monolingual anchoring may be further explored for related low-resource languages, such as Kazakh, Kyrgyz, and Uzbek.

Original languageEnglish
Article number736
JournalJournal of King Saud University - Computer and Information Sciences
Volume38
Issue number7
DOIs
Publication statusPublished - Sept 2026

Keywords

  • Cross-lingual semantic alignment
  • Low-resource language
  • Monolingual anchoring
  • Sentence embeddings
  • Uyghur

Fingerprint

Dive into the research topics of 'Monolingual anchoring for low-resource cross-lingual semantic alignment: A case study on uyghur'. Together they form a unique fingerprint.

Cite this