跳到主要导航 跳到搜索 跳到主要内容

TIGSEN: Building a Low-Resource Dataset and Benchmarking for Tigrigna Sentiment Analysis with Cross-Lingual Transfer Learning Approaches

  • Hagos Gebremedhin Gebremeskel*
  • , Chong Feng*
  • , Asefa Mebrahtu Abera
  • *此作品的通讯作者
  • Beijing Institute of Technology
  • Mekelle University
  • Aksum University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Sentiment analysis in low-resource languages faces several challenges. This paper addresses the challenges of extremely low-resource and methodologically understudied sentiment analysis for Tigrigna, an official language spoken in Eritrea and the Tigray region of Ethiopia. We introduce TIGSEN, the first large-scale, multi-domain benchmark dataset for Tigrigna sentiment analysis, comprising 68,596 annotated text samples from official social media, news, and review forums. Created through a rigorous native-speakers annotation protocol, the dataset is designed to enable robust model training and evaluation. To prove its utility and to establish a strong, reproducible baseline for the community, we propose and evaluate a systematic cross-transfer learning framework. This methodology deliberately leverages annotated data from high-resource to linguistically related languages, English and Amharic, to overcome the limitations of Tigrigna’s small data pool. Our experiments show that models fine-tuned directly on TIGSEN achieved competitive performance. At the same time, the proposed cross-transfer framework yields a significant performance gain, achieving an accuracy of 87.6% over a strong multilingual baseline. This result validates TIGSEN as a learnable and challenging benchmark and provides a practical blueprint for resource amplification in low-resource settings. We publicly release the TIGSEN dataset, annotation guidelines, and benchmarking code to serve as a foundational resource and a catalyst for future LLM research in Tigrigna NLP.

源语言英语
主期刊名Machine Translation - 21st China Conference, CCMT 2025, Proceedings
编辑Jin'an Xu, Zhaopeng Tu, Kehai Chen, Yuhang Guo
出版商Springer Science and Business Media Deutschland GmbH
1-17
页数17
ISBN(印刷版)9789819201983
DOI
出版状态已出版 - 2026
已对外发布
活动21st China Conference on Machine Translation, CCMT 2025 - Lanzhou, 中国
期限: 26 9月 202528 9月 2025

丛书

姓名Communications in Computer and Information Science
2906 CCIS
ISSN(印刷版)1865-0929
ISSN(电子版)1865-0937

会议

会议21st China Conference on Machine Translation, CCMT 2025
国家/地区中国
Lanzhou
时期26/09/2528/09/25

学术指纹

探究 'TIGSEN: Building a Low-Resource Dataset and Benchmarking for Tigrigna Sentiment Analysis with Cross-Lingual Transfer Learning Approaches' 的科研主题。它们共同构成独一无二的学术指纹。

引用此