TY - JOUR
T1 - Central Dynamic Transformer for Hyperspectral Image Classification
AU - Zhang, Zhibin
AU - Liu, Huan
AU - Zheng, Zhiyang
AU - Qu, Wenyu
AU - Li, Wei
N1 - Publisher Copyright:
© 2008-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - In hyperspectral image (HSI) classification, patch-based methods have evolved from treating all pixels equally to acknowledging the importance of central pixels whose labels determine classification outcomes. However, existing central attention approaches suffer from two limitations: they rely on static linear transformations that apply fixed weights uniformly across inputs, and they focus solely on the central pixel while neglecting other same-object pixels, leading to incomplete representations. To address these limitations, we propose the Central Dynamic Transformer (CDTrans) through two innovations. First, the Central Dynamic Linear (CDLinear) module replaces static projections with spatially adaptive transformations, modulating key and value representations based on global context and central pixel characteristics. Second, Central Dynamic Attention (CDAttn) module employs a two-stage mechanism: identifying the K most semantically similar pixels through cosine similarity in the CDLinear-transformed feature space, then performing attention where selected pixels attend to all pixels. This enables discriminative object-level feature extraction. Experiments on four datasets demonstrate that CDTrans achieves state-of-the-art performance, with ablation studies validating the significant contributions of both CDLinear and CDAttn.
AB - In hyperspectral image (HSI) classification, patch-based methods have evolved from treating all pixels equally to acknowledging the importance of central pixels whose labels determine classification outcomes. However, existing central attention approaches suffer from two limitations: they rely on static linear transformations that apply fixed weights uniformly across inputs, and they focus solely on the central pixel while neglecting other same-object pixels, leading to incomplete representations. To address these limitations, we propose the Central Dynamic Transformer (CDTrans) through two innovations. First, the Central Dynamic Linear (CDLinear) module replaces static projections with spatially adaptive transformations, modulating key and value representations based on global context and central pixel characteristics. Second, Central Dynamic Attention (CDAttn) module employs a two-stage mechanism: identifying the K most semantically similar pixels through cosine similarity in the CDLinear-transformed feature space, then performing attention where selected pixels attend to all pixels. This enables discriminative object-level feature extraction. Experiments on four datasets demonstrate that CDTrans achieves state-of-the-art performance, with ablation studies validating the significant contributions of both CDLinear and CDAttn.
KW - Central dynamic attention (CDAttn)
KW - central dynamic linear (CDLinear)
KW - dynamic token selection
KW - hyperspectral image (HSI)
UR - https://www.scopus.com/pages/publications/105039140635
U2 - 10.1109/JSTARS.2026.3693282
DO - 10.1109/JSTARS.2026.3693282
M3 - Article
AN - SCOPUS:105039140635
SN - 1939-1404
VL - 19
SP - 17363
EP - 17378
JO - IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing
JF - IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing
ER -