TY - JOUR
T1 - ESTR-CoT
T2 - Towards explainable and accurate event stream based scene text recognition with chain-of-thought reasoning
AU - Wang, Xiao
AU - Jiang, Jingtao
AU - Chen, Qiang
AU - Chen, Lan
AU - Zhu, Lin
AU - Wang, Yaowei
AU - Tian, Yonghong
AU - Tang, Jin
N1 - Publisher Copyright:
© 2026 Elsevier B.V.
PY - 2026/10/28
Y1 - 2026/10/28
N2 - Event stream based scene text recognition is a newly arising research topic in recent years which performs better than the widely used RGB cameras in extremely challenging scenarios, especially in low illumination, and fast motion. Existing works either adopt an end-to-end encoder-decoder framework or large language models for enhanced recognition; however, they are still limited by the challenges of insufficient interpretability and weak contextual logical reasoning. In this work, we propose a novel chain-of-thought reasoning based event stream scene text recognition framework, termed ESTR-CoT. Specifically, we first adopt the vision encoder EVA-CLIP (ViT-G/14) to transform the input event stream into tokens and utilize a Llama tokenizer to encode the given generation prompt. A Q-former is used to align the vision token to the pre-trained large language model Vicuna-7B and output both the answer and chain-of-thought (CoT) reasoning process simultaneously. Our framework can be optimized using supervised fine-tuning in an end-to-end manner. In addition, we also propose a large-scale CoT dataset to train our framework via a three stage processing (i.e., generation, polish, and expert verification). This dataset provides a solid data foundation for the development of subsequent reasoning-based large models. Extensive experiments on three event stream STR benchmark datasets (i.e., EventSTR, WordArt*, IC15*) fully validated the effectiveness and interpretability of our proposed framework. The source code and pre-trained models will be released on https://github.com/Event-AHU/ESTR-CoT.
AB - Event stream based scene text recognition is a newly arising research topic in recent years which performs better than the widely used RGB cameras in extremely challenging scenarios, especially in low illumination, and fast motion. Existing works either adopt an end-to-end encoder-decoder framework or large language models for enhanced recognition; however, they are still limited by the challenges of insufficient interpretability and weak contextual logical reasoning. In this work, we propose a novel chain-of-thought reasoning based event stream scene text recognition framework, termed ESTR-CoT. Specifically, we first adopt the vision encoder EVA-CLIP (ViT-G/14) to transform the input event stream into tokens and utilize a Llama tokenizer to encode the given generation prompt. A Q-former is used to align the vision token to the pre-trained large language model Vicuna-7B and output both the answer and chain-of-thought (CoT) reasoning process simultaneously. Our framework can be optimized using supervised fine-tuning in an end-to-end manner. In addition, we also propose a large-scale CoT dataset to train our framework via a three stage processing (i.e., generation, polish, and expert verification). This dataset provides a solid data foundation for the development of subsequent reasoning-based large models. Extensive experiments on three event stream STR benchmark datasets (i.e., EventSTR, WordArt*, IC15*) fully validated the effectiveness and interpretability of our proposed framework. The source code and pre-trained models will be released on https://github.com/Event-AHU/ESTR-CoT.
KW - Chain-of-thought reasoning
KW - Event camera
KW - Explainable artificial intelligence
KW - Large language models
KW - Scene text recognition
UR - https://www.scopus.com/pages/publications/105043158765
U2 - 10.1016/j.neucom.2026.134311
DO - 10.1016/j.neucom.2026.134311
M3 - Article
AN - SCOPUS:105043158765
SN - 0925-2312
VL - 699
JO - Neurocomputing
JF - Neurocomputing
M1 - 134311
ER -