TY - GEN
T1 - Multimodal Image Registration via Contrastive Learning and Multi-Scale Progressive Deformation Estimation
AU - Shen, Hengyu
AU - Chen, Jiajing
AU - Zhou, Zhiqiang
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Multimodal images can provide richer scene information. However, due to differences in imaging mechanisms, infrared and visible images often exhibit significant modality differences and spatial misalignments, which pose challenges for registration and subsequent fusion tasks. To address this issue, this paper proposes a multimodal image registration method based on contrastive learning and multiscale progressive deformation estimation-CMPE. The method first introduces a Contrastive Learning Module (CLM) to extract cross-modal shared semantic features, significantly reducing the modality gap between infrared and visible images. Subsequently, a multiscale progressive registration framework based on an encoder-decoder structure is designed, and a Global-Local Attention Module (GLAM) is employed at each scale to adaptively select the features, which are then used to predict the deformation field at each scale. The multiscale predicted deformation fields are adaptively weighted and smoothed through a Dynamic Field Fusion Module (DFFM) and vector field integration, ensuring the continuity of the output deformation. Extensive experiments demonstrate that CMPE outperforms existing methods in both qualitative and quantitative evaluations.
AB - Multimodal images can provide richer scene information. However, due to differences in imaging mechanisms, infrared and visible images often exhibit significant modality differences and spatial misalignments, which pose challenges for registration and subsequent fusion tasks. To address this issue, this paper proposes a multimodal image registration method based on contrastive learning and multiscale progressive deformation estimation-CMPE. The method first introduces a Contrastive Learning Module (CLM) to extract cross-modal shared semantic features, significantly reducing the modality gap between infrared and visible images. Subsequently, a multiscale progressive registration framework based on an encoder-decoder structure is designed, and a Global-Local Attention Module (GLAM) is employed at each scale to adaptively select the features, which are then used to predict the deformation field at each scale. The multiscale predicted deformation fields are adaptively weighted and smoothed through a Dynamic Field Fusion Module (DFFM) and vector field integration, ensuring the continuity of the output deformation. Extensive experiments demonstrate that CMPE outperforms existing methods in both qualitative and quantitative evaluations.
KW - contrastive learning
KW - Multimodal image registration
KW - progressive deformation estimation
UR - https://www.scopus.com/pages/publications/105043924084
U2 - 10.1109/CCDC69976.2026.11560196
DO - 10.1109/CCDC69976.2026.11560196
M3 - Conference contribution
AN - SCOPUS:105043924084
T3 - 38th Chinese Control and Decision Conference, CCDC 2026
SP - 2558
EP - 2565
BT - 38th Chinese Control and Decision Conference, CCDC 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 38th Chinese Control and Decision Conference, CCDC 2026
Y2 - 15 May 2026 through 18 May 2026
ER -