TY - JOUR
T1 - Shape Activated CAM Learning for Weakly Supervised Remote Sensing Semantic Segmentation
AU - Chen, He
AU - Dong, Mingyue
AU - Yue, Linwei
AU - Zheng, Xianwei
AU - Li, Jun
AU - Gong, Jianya
N1 - Publisher Copyright:
© 1980-2012 IEEE.
PY - 2025
Y1 - 2025
N2 - Class activation map (CAM)-based weakly supervised semantic segmentation (WSSS) of remote sensing (RS) images has attracted extensive research interests for its potential in reducing annotation cost. However, challenged by unconstrained activation issue, existing methods struggle to delineate object boundaries clearly, making them particularly difficult to separate multiple densely packed objects, which are common in RS images. By conducting an in-depth analysis of RS image characteristics, we observed a strong correlation between object shapes and their semantics. Inspired by this finding, we propose an intrinsic shape activation network (ISANet) to learn the category-relevant shape priors as geometry constraints for target-focused region activation in WSSS of RS images. The key idea is to distill the intrinsic shape priors from the hybrid features that are deterministic in classification. Specifically, we adopt a dual-branch architecture to decouple the learning of shape and texture features and leverage a shape awareness alignment module (SAM) to generate boundary-clear CAMs for computing pseudo-labels. In this way, CAMs are generated with perception of target shapes, which increases the completeness of activation regions and alleviates the ultrarange responses. Extensive experiments demonstrate the superiority of our method in delineating densely packed objects with clear contours, which is especially beneficial for separating multiple targets in RS images. Our method improves the mean intersection over union (mIoU) of the state-of-the-art method by 7.9% and 3.3% on the NWPU VHR-10 and iSAID dataset, respectively.
AB - Class activation map (CAM)-based weakly supervised semantic segmentation (WSSS) of remote sensing (RS) images has attracted extensive research interests for its potential in reducing annotation cost. However, challenged by unconstrained activation issue, existing methods struggle to delineate object boundaries clearly, making them particularly difficult to separate multiple densely packed objects, which are common in RS images. By conducting an in-depth analysis of RS image characteristics, we observed a strong correlation between object shapes and their semantics. Inspired by this finding, we propose an intrinsic shape activation network (ISANet) to learn the category-relevant shape priors as geometry constraints for target-focused region activation in WSSS of RS images. The key idea is to distill the intrinsic shape priors from the hybrid features that are deterministic in classification. Specifically, we adopt a dual-branch architecture to decouple the learning of shape and texture features and leverage a shape awareness alignment module (SAM) to generate boundary-clear CAMs for computing pseudo-labels. In this way, CAMs are generated with perception of target shapes, which increases the completeness of activation regions and alleviates the ultrarange responses. Extensive experiments demonstrate the superiority of our method in delineating densely packed objects with clear contours, which is especially beneficial for separating multiple targets in RS images. Our method improves the mean intersection over union (mIoU) of the state-of-the-art method by 7.9% and 3.3% on the NWPU VHR-10 and iSAID dataset, respectively.
KW - Class activation map (CAM)
KW - remote sensing (RS)
KW - shape prior
KW - weakly supervised semantic segmentation (WSSS)
UR - https://www.scopus.com/pages/publications/105007873453
U2 - 10.1109/TGRS.2025.3578466
DO - 10.1109/TGRS.2025.3578466
M3 - Article
AN - SCOPUS:105007873453
SN - 0196-2892
VL - 63
JO - IEEE Transactions on Geoscience and Remote Sensing
JF - IEEE Transactions on Geoscience and Remote Sensing
M1 - 5627516
ER -