TY - GEN
T1 - TIGSEN
T2 - 21st China Conference on Machine Translation, CCMT 2025
AU - Gebremeskel, Hagos Gebremedhin
AU - Feng, Chong
AU - Abera, Asefa Mebrahtu
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - Sentiment analysis in low-resource languages faces several challenges. This paper addresses the challenges of extremely low-resource and methodologically understudied sentiment analysis for Tigrigna, an official language spoken in Eritrea and the Tigray region of Ethiopia. We introduce TIGSEN, the first large-scale, multi-domain benchmark dataset for Tigrigna sentiment analysis, comprising 68,596 annotated text samples from official social media, news, and review forums. Created through a rigorous native-speakers annotation protocol, the dataset is designed to enable robust model training and evaluation. To prove its utility and to establish a strong, reproducible baseline for the community, we propose and evaluate a systematic cross-transfer learning framework. This methodology deliberately leverages annotated data from high-resource to linguistically related languages, English and Amharic, to overcome the limitations of Tigrigna’s small data pool. Our experiments show that models fine-tuned directly on TIGSEN achieved competitive performance. At the same time, the proposed cross-transfer framework yields a significant performance gain, achieving an accuracy of 87.6% over a strong multilingual baseline. This result validates TIGSEN as a learnable and challenging benchmark and provides a practical blueprint for resource amplification in low-resource settings. We publicly release the TIGSEN dataset, annotation guidelines, and benchmarking code to serve as a foundational resource and a catalyst for future LLM research in Tigrigna NLP.
AB - Sentiment analysis in low-resource languages faces several challenges. This paper addresses the challenges of extremely low-resource and methodologically understudied sentiment analysis for Tigrigna, an official language spoken in Eritrea and the Tigray region of Ethiopia. We introduce TIGSEN, the first large-scale, multi-domain benchmark dataset for Tigrigna sentiment analysis, comprising 68,596 annotated text samples from official social media, news, and review forums. Created through a rigorous native-speakers annotation protocol, the dataset is designed to enable robust model training and evaluation. To prove its utility and to establish a strong, reproducible baseline for the community, we propose and evaluate a systematic cross-transfer learning framework. This methodology deliberately leverages annotated data from high-resource to linguistically related languages, English and Amharic, to overcome the limitations of Tigrigna’s small data pool. Our experiments show that models fine-tuned directly on TIGSEN achieved competitive performance. At the same time, the proposed cross-transfer framework yields a significant performance gain, achieving an accuracy of 87.6% over a strong multilingual baseline. This result validates TIGSEN as a learnable and challenging benchmark and provides a practical blueprint for resource amplification in low-resource settings. We publicly release the TIGSEN dataset, annotation guidelines, and benchmarking code to serve as a foundational resource and a catalyst for future LLM research in Tigrigna NLP.
KW - Cross-Lingual Transfer
KW - Low-Resource
KW - Sentiment Analysis
KW - Tigrigna
KW - TIGSEN
UR - https://www.scopus.com/pages/publications/105043992070
U2 - 10.1007/978-981-92-0199-0_1
DO - 10.1007/978-981-92-0199-0_1
M3 - Conference contribution
AN - SCOPUS:105043992070
SN - 9789819201983
T3 - Communications in Computer and Information Science
SP - 1
EP - 17
BT - Machine Translation - 21st China Conference, CCMT 2025, Proceedings
A2 - Xu, Jin'an
A2 - Tu, Zhaopeng
A2 - Chen, Kehai
A2 - Guo, Yuhang
PB - Springer Science and Business Media Deutschland GmbH
Y2 - 26 September 2025 through 28 September 2025
ER -