TY - JOUR
T1 - STRAViT
T2 - A Spatial-Temporal Feature-Reshaped Agent Vision Transformer for EEG-Based Emotion Recognition
AU - Cui, Xinyu
AU - Li, Xiaowei
AU - Zhu, Jing
AU - Hu, Bin
N1 - Publisher Copyright:
© 1963-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Emotion recognition from electroencephalogram (EEG) signals has attracted increasing interest due to its applications in affective computing and brain-computer interaction. However, effectively integrating spatial-temporal features while attending to emotionally relevant brain regions remains a significant challenge. In this article, we propose a spatial-temporal feature-reshaped agent vision transformer (STRAViT) network, which fuses differential entropy (DE) and functional connectivity matrices through a dual-branch attention-based feature reshaping module (FRM), enabling refined feature integration. In addition, we introduce an agent vision transformer (AViT) that utilizes learnable agent tokens to capture global-local dependencies within EEG representations efficiently. Extensive experiments conducted on the SEED and SEED-IV datasets demonstrate that the STRAViT achieves classification accuracies of 98.7% and 87.2% on SEED, 95.5% and 72.5% on SEED-IV, for discrete emotion recognition, under subject-dependent and subject-independent strategies. The proposed network not only effectively integrates spatial-temporal features of EEG signals but also enhances the modeling of inter-regional dependencies through agent attention. Comprehensive tests confirm that the STRAViT yields excellent performance on emotion recognition tasks.
AB - Emotion recognition from electroencephalogram (EEG) signals has attracted increasing interest due to its applications in affective computing and brain-computer interaction. However, effectively integrating spatial-temporal features while attending to emotionally relevant brain regions remains a significant challenge. In this article, we propose a spatial-temporal feature-reshaped agent vision transformer (STRAViT) network, which fuses differential entropy (DE) and functional connectivity matrices through a dual-branch attention-based feature reshaping module (FRM), enabling refined feature integration. In addition, we introduce an agent vision transformer (AViT) that utilizes learnable agent tokens to capture global-local dependencies within EEG representations efficiently. Extensive experiments conducted on the SEED and SEED-IV datasets demonstrate that the STRAViT achieves classification accuracies of 98.7% and 87.2% on SEED, 95.5% and 72.5% on SEED-IV, for discrete emotion recognition, under subject-dependent and subject-independent strategies. The proposed network not only effectively integrates spatial-temporal features of EEG signals but also enhances the modeling of inter-regional dependencies through agent attention. Comprehensive tests confirm that the STRAViT yields excellent performance on emotion recognition tasks.
KW - Agent vision transformer (AViT)
KW - attention
KW - electroencephalogram (EEG) emotion recognition
KW - feature fusion
UR - https://www.scopus.com/pages/publications/105043504014
U2 - 10.1109/TIM.2026.3701192
DO - 10.1109/TIM.2026.3701192
M3 - Article
AN - SCOPUS:105043504014
SN - 0018-9456
VL - 75
JO - IEEE Transactions on Instrumentation and Measurement
JF - IEEE Transactions on Instrumentation and Measurement
M1 - 4009815
ER -