TY - JOUR
T1 - LiteNER
T2 - A novel lightweight method for long text named entity recognition
AU - Chen, Yelin
AU - Zhang, Huaping
AU - Yan, Ruohao
AU - Zhu, Jihong
AU - Hamdulla, Askar
N1 - Publisher Copyright:
© 2026 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
PY - 2027/1
Y1 - 2027/1
N2 - Despite extensive research on Named Entity Recognition (NER), accurate and efficient extraction from long texts, such as scholar homepages, remains an underexplored challenge. When applied to long texts, existing approaches often become computationally expensive and less accurate, limiting their scalability in practice. In this paper, we introduce LiteNER, a lightweight framework for long-text NER. It incorporates an Anchor-Gathering Attention (AGA) mechanism to inject global contextual cues at reduced interaction cost, an Adaptive Differentiable Token Filtering (ADTF) strategy to discard non-essential tokens while preserving boundary-relevant information, and a Staircase Dual-Axis Span Interaction (SDA-SI) module to reduce redundant span interactions during entity extraction. Comprehensive empirical evaluations on three long-text NER datasets show that LiteNER outperforms the strongest baseline by up to 2.46% and 1.07% in F1 score on Scholar-XL and SciREX, while maintaining comparable performance on Profiling-07. It further supports input sequences over 11.6 times longer and achieves up to 2.3× faster inference speed, thereby underscoring its practical efficacy for long-form text NER.
AB - Despite extensive research on Named Entity Recognition (NER), accurate and efficient extraction from long texts, such as scholar homepages, remains an underexplored challenge. When applied to long texts, existing approaches often become computationally expensive and less accurate, limiting their scalability in practice. In this paper, we introduce LiteNER, a lightweight framework for long-text NER. It incorporates an Anchor-Gathering Attention (AGA) mechanism to inject global contextual cues at reduced interaction cost, an Adaptive Differentiable Token Filtering (ADTF) strategy to discard non-essential tokens while preserving boundary-relevant information, and a Staircase Dual-Axis Span Interaction (SDA-SI) module to reduce redundant span interactions during entity extraction. Comprehensive empirical evaluations on three long-text NER datasets show that LiteNER outperforms the strongest baseline by up to 2.46% and 1.07% in F1 score on Scholar-XL and SciREX, while maintaining comparable performance on Profiling-07. It further supports input sequences over 11.6 times longer and achieves up to 2.3× faster inference speed, thereby underscoring its practical efficacy for long-form text NER.
KW - Deep learning
KW - Long-sequence encoding
KW - Long-text NER
KW - Span interaction modeling
KW - Token filtering
UR - https://www.scopus.com/pages/publications/105043409726
U2 - 10.1016/j.ipm.2026.104981
DO - 10.1016/j.ipm.2026.104981
M3 - Article
AN - SCOPUS:105043409726
SN - 0306-4573
VL - 64
JO - Information Processing and Management
JF - Information Processing and Management
IS - 1
M1 - 104981
ER -