TY - JOUR
T1 - Monolingual anchoring for low-resource cross-lingual semantic alignment
T2 - A case study on uyghur
AU - Yan, Ruohao
AU - Zhang, Huaping
AU - Niu, Yuwen
AU - Zhu, Jihong
AU - Hamdulla, Askar
N1 - Publisher Copyright:
© The Author(s) 2026.
PY - 2026/9
Y1 - 2026/9
N2 - Low-resource languages remain challenging for cross-lingual semantic alignment because of limited parallel corpora. In addition, conventional symmetric alignment may distort the semantic space of a high-resource language through noisy low-resource updates. To address this issue, we propose Monolingual Anchoring for Cross-Lingual Semantic Alignment (MACA), focusing on Uyghur as a low-resource case study. MACA follows an asymmetric paradigm that treats the high-resource language as a fixed semantic anchor and transfers its semantic structure to the Uyghur side. The method consists of three components: (1) Anchored Embedding Initialization for newly introduced Uyghur subwords, (2) Cross-Lingual Neighborhood Anchoring for structural alignment between Uyghur and the anchor language, and (3) Monolingual Structure Anchoring for improving the internal semantic organization of Uyghur representations. Experiments centered on Uyghur-Chinese show that MACA outperforms LaBSE, the strongest off-the-shelf multilingual baseline in our comparison, by 7.67 points on cross-lingual STS. In an exploratory Uyghur-English zero-shot setting, MACA also surpasses LaBSE by 2.45 points without using Uyghur-English training data. These results provide evidence for the effectiveness of MACA in the evaluated Uyghur setting and suggest that monolingual anchoring may be further explored for related low-resource languages, such as Kazakh, Kyrgyz, and Uzbek.
AB - Low-resource languages remain challenging for cross-lingual semantic alignment because of limited parallel corpora. In addition, conventional symmetric alignment may distort the semantic space of a high-resource language through noisy low-resource updates. To address this issue, we propose Monolingual Anchoring for Cross-Lingual Semantic Alignment (MACA), focusing on Uyghur as a low-resource case study. MACA follows an asymmetric paradigm that treats the high-resource language as a fixed semantic anchor and transfers its semantic structure to the Uyghur side. The method consists of three components: (1) Anchored Embedding Initialization for newly introduced Uyghur subwords, (2) Cross-Lingual Neighborhood Anchoring for structural alignment between Uyghur and the anchor language, and (3) Monolingual Structure Anchoring for improving the internal semantic organization of Uyghur representations. Experiments centered on Uyghur-Chinese show that MACA outperforms LaBSE, the strongest off-the-shelf multilingual baseline in our comparison, by 7.67 points on cross-lingual STS. In an exploratory Uyghur-English zero-shot setting, MACA also surpasses LaBSE by 2.45 points without using Uyghur-English training data. These results provide evidence for the effectiveness of MACA in the evaluated Uyghur setting and suggest that monolingual anchoring may be further explored for related low-resource languages, such as Kazakh, Kyrgyz, and Uzbek.
KW - Cross-lingual semantic alignment
KW - Low-resource language
KW - Monolingual anchoring
KW - Sentence embeddings
KW - Uyghur
UR - https://www.scopus.com/pages/publications/105047853953
U2 - 10.1007/s44443-026-01019-4
DO - 10.1007/s44443-026-01019-4
M3 - Article
AN - SCOPUS:105047853953
SN - 1319-1578
VL - 38
JO - Journal of King Saud University - Computer and Information Sciences
JF - Journal of King Saud University - Computer and Information Sciences
IS - 7
M1 - 736
ER -