Abstract
Despite extensive research on Named Entity Recognition (NER), accurate and efficient extraction from long texts, such as scholar homepages, remains an underexplored challenge. When applied to long texts, existing approaches often become computationally expensive and less accurate, limiting their scalability in practice. In this paper, we introduce LiteNER, a lightweight framework for long-text NER. It incorporates an Anchor-Gathering Attention (AGA) mechanism to inject global contextual cues at reduced interaction cost, an Adaptive Differentiable Token Filtering (ADTF) strategy to discard non-essential tokens while preserving boundary-relevant information, and a Staircase Dual-Axis Span Interaction (SDA-SI) module to reduce redundant span interactions during entity extraction. Comprehensive empirical evaluations on three long-text NER datasets show that LiteNER outperforms the strongest baseline by up to 2.46% and 1.07% in F1 score on Scholar-XL and SciREX, while maintaining comparable performance on Profiling-07. It further supports input sequences over 11.6 times longer and achieves up to 2.3× faster inference speed, thereby underscoring its practical efficacy for long-form text NER.
| Original language | English |
|---|---|
| Article number | 104981 |
| Journal | Information Processing and Management |
| Volume | 64 |
| Issue number | 1 |
| DOIs | |
| Publication status | Published - Jan 2027 |
| Externally published | Yes |
Keywords
- Deep learning
- Long-sequence encoding
- Long-text NER
- Span interaction modeling
- Token filtering
Fingerprint
Dive into the research topics of 'LiteNER: A novel lightweight method for long text named entity recognition'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver