TY - JOUR
T1 - Structure-Aware Mask and Detail-Augmented Transformer Network for Remote Hyperspectral Image Classification
AU - Wang, Zhaorui
AU - Huang, Bin
AU - Qin, Tong
N1 - Publisher Copyright:
© 1980-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Transformer networks exhibit strong capabilities in modeling global spatial features and have achieved remarkable success in hyperspectral image (HSI) classification. However, Transformer-based methods also face several challenges, including a substantial increase in model parameters, strong dependence on abundant labeled samples, and difficulty effectively fusing spatial and spectral information. To address these issues, we propose a Structure-Aware Mask and Detail-Augmented Transformer (SAM-DAT) network. Specifically, a Spatial-Spectral Hybrid Perceptual Masking (SSHPM) strategy is proposed to highlight boundary regions based on pixel-level similarity analysis, thereby mitigating the loss of fine-grained information and noise introduction caused by conventional masking methods. A Local Spatial-Spectral Feature Extraction (LSSFE) module is designed by integrating two-dimension (2-D) and three-dimension (3-D) convolutions to jointly model spatial and spectral features at multiple scales. A Global Fine-Grained Augmentation (GFGA) Transformer is introduced to fuse local details and global dependencies through comprehensive feature encoding and a novel attention mechanism. Finally, a Decoupled Output Head (DOH) module is designed to simultaneously perform image reconstruction and classification, promoting consistency between discriminative and reconstructive feature learning. Extensive experiments on four benchmark HSI classification datasets demonstrate that the proposed SAM-DAT significantly outperforms existing state-of-the-art methods.
AB - Transformer networks exhibit strong capabilities in modeling global spatial features and have achieved remarkable success in hyperspectral image (HSI) classification. However, Transformer-based methods also face several challenges, including a substantial increase in model parameters, strong dependence on abundant labeled samples, and difficulty effectively fusing spatial and spectral information. To address these issues, we propose a Structure-Aware Mask and Detail-Augmented Transformer (SAM-DAT) network. Specifically, a Spatial-Spectral Hybrid Perceptual Masking (SSHPM) strategy is proposed to highlight boundary regions based on pixel-level similarity analysis, thereby mitigating the loss of fine-grained information and noise introduction caused by conventional masking methods. A Local Spatial-Spectral Feature Extraction (LSSFE) module is designed by integrating two-dimension (2-D) and three-dimension (3-D) convolutions to jointly model spatial and spectral features at multiple scales. A Global Fine-Grained Augmentation (GFGA) Transformer is introduced to fuse local details and global dependencies through comprehensive feature encoding and a novel attention mechanism. Finally, a Decoupled Output Head (DOH) module is designed to simultaneously perform image reconstruction and classification, promoting consistency between discriminative and reconstructive feature learning. Extensive experiments on four benchmark HSI classification datasets demonstrate that the proposed SAM-DAT significantly outperforms existing state-of-the-art methods.
KW - Deep learning
KW - Hyperspectral image
KW - Image classification
KW - Transformer
UR - https://www.scopus.com/pages/publications/105040159260
U2 - 10.1109/TGRS.2026.3697151
DO - 10.1109/TGRS.2026.3697151
M3 - Article
AN - SCOPUS:105040159260
SN - 0196-2892
JO - IEEE Transactions on Geoscience and Remote Sensing
JF - IEEE Transactions on Geoscience and Remote Sensing
ER -