Skip to main navigation Skip to search Skip to main content

TIGSEN: Building a Low-Resource Dataset and Benchmarking for Tigrigna Sentiment Analysis with Cross-Lingual Transfer Learning Approaches

  • Hagos Gebremedhin Gebremeskel*
  • , Chong Feng*
  • , Asefa Mebrahtu Abera
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • Mekelle University
  • Aksum University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Sentiment analysis in low-resource languages faces several challenges. This paper addresses the challenges of extremely low-resource and methodologically understudied sentiment analysis for Tigrigna, an official language spoken in Eritrea and the Tigray region of Ethiopia. We introduce TIGSEN, the first large-scale, multi-domain benchmark dataset for Tigrigna sentiment analysis, comprising 68,596 annotated text samples from official social media, news, and review forums. Created through a rigorous native-speakers annotation protocol, the dataset is designed to enable robust model training and evaluation. To prove its utility and to establish a strong, reproducible baseline for the community, we propose and evaluate a systematic cross-transfer learning framework. This methodology deliberately leverages annotated data from high-resource to linguistically related languages, English and Amharic, to overcome the limitations of Tigrigna’s small data pool. Our experiments show that models fine-tuned directly on TIGSEN achieved competitive performance. At the same time, the proposed cross-transfer framework yields a significant performance gain, achieving an accuracy of 87.6% over a strong multilingual baseline. This result validates TIGSEN as a learnable and challenging benchmark and provides a practical blueprint for resource amplification in low-resource settings. We publicly release the TIGSEN dataset, annotation guidelines, and benchmarking code to serve as a foundational resource and a catalyst for future LLM research in Tigrigna NLP.

Original languageEnglish
Title of host publicationMachine Translation - 21st China Conference, CCMT 2025, Proceedings
EditorsJin'an Xu, Zhaopeng Tu, Kehai Chen, Yuhang Guo
PublisherSpringer Science and Business Media Deutschland GmbH
Pages1-17
Number of pages17
ISBN (Print)9789819201983
DOIs
Publication statusPublished - 2026
Externally publishedYes
Event21st China Conference on Machine Translation, CCMT 2025 - Lanzhou, China
Duration: 26 Sept 202528 Sept 2025

Publication series

NameCommunications in Computer and Information Science
Volume2906 CCIS
ISSN (Print)1865-0929
ISSN (Electronic)1865-0937

Conference

Conference21st China Conference on Machine Translation, CCMT 2025
Country/TerritoryChina
CityLanzhou
Period26/09/2528/09/25

Keywords

  • Cross-Lingual Transfer
  • Low-Resource
  • Sentiment Analysis
  • Tigrigna
  • TIGSEN

Fingerprint

Dive into the research topics of 'TIGSEN: Building a Low-Resource Dataset and Benchmarking for Tigrigna Sentiment Analysis with Cross-Lingual Transfer Learning Approaches'. Together they form a unique fingerprint.

Cite this