Abstract
In hyperspectral image (HSI) classification, patch-based methods have evolved from treating all pixels equally to acknowledging the importance of central pixels whose labels determine classification outcomes. However, existing central attention approaches suffer from two limitations: they rely on static linear transformations that apply fixed weights uniformly across inputs, and they focus solely on the central pixel while neglecting other same-object pixels, leading to incomplete representations. To address these limitations, we propose the Central Dynamic Transformer (CDTrans) through two innovations. First, the Central Dynamic Linear (CDLinear) module replaces static projections with spatially adaptive transformations, modulating key and value representations based on global context and central pixel characteristics. Second, Central Dynamic Attention (CDAttn) module employs a two-stage mechanism: identifying the K most semantically similar pixels through cosine similarity in the CDLinear-transformed feature space, then performing attention where selected pixels attend to all pixels. This enables discriminative object-level feature extraction. Experiments on four datasets demonstrate that CDTrans achieves state-of-the-art performance, with ablation studies validating the significant contributions of both CDLinear and CDAttn.
| Original language | English |
|---|---|
| Pages (from-to) | 17363-17378 |
| Number of pages | 16 |
| Journal | IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing |
| Volume | 19 |
| DOIs | |
| Publication status | Published - 2026 |
Keywords
- Central dynamic attention (CDAttn)
- central dynamic linear (CDLinear)
- dynamic token selection
- hyperspectral image (HSI)
Fingerprint
Dive into the research topics of 'Central Dynamic Transformer for Hyperspectral Image Classification'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver