Skip to main navigation Skip to search Skip to main content

基于二阶注意力混合尺度特征补偿的多模态图像融合方法

Translated title of the contribution: Multimodal image fusion method based on second-order attention and mixed-scale feature compensation
  • Beijing Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Objective In the field of multimodal information processing, infrared and visible image fusion technology has emerged as a research hotspot in computer vision. This technology integrates the rich texture and color details of visible images with the superior thermal radiation penetration capability of infrared images. Traditional methods suffer from limitations such as high complexity in fusion rule design and stringent application constraints. Although existing deep learning-based approaches have achieved remarkable results, they still face several challenges: target structures and texture details are prone to loss during feature extraction; feature fusion mechanisms struggle to capture high-order dependencies; and multi-scale fusion strategies overlook the cross-modal mixed-scale complementarity. To address these issues, we proposes a multimodal image fusion method based on second-order attention and mixed-scale feature compensation (SAMFuse). Methods The SAMFuse method (Fig.1) employs a dual-branch input architecture. First, it extracts their initial features via convolution operations upon feeding a pair of infrared and visible images; Second, the Weighted-Residual Feature Extraction Module (Fig.3) leverages Depthwise Separable Convolution (Fig.2) with varying strides, generating multi-scale representations by fusing features from the main path and the weighted residual path, where high-scale features serve to preserve detailed information while low-scale features capture contextual semantic content; Third, intra-scale and cross-scale features derived from different modalities are integrated through the Mixed-Scale Second-Order Attention Fusion Module (Fig.5) to yield multi-scale feature maps; Finally, these multi-scale features are concatenated and fed into the Multi-Scale Image Reconstruction Module (Fig.4), which produces the final fused image. To concurrently retain the thermal radiation information of infrared images and the fine-grained texture details of visible images, we have devised a multi-scale adaptive loss function. Results and Discussions Qualitative and quantitative evaluations are conducted on two public datasets, namely TNO and M3FD. Scene 1-Scene 3 in Fig.8 present the visualization results on the TNO dataset, demonstrating that the proposed method can not only accurately preserve the pixel intensity of infrared highlight targets and visible light bright regions, but also retain infrared regions with drastic intensity variations and visible light texture details. As shown in Tab.1, the method achieves the highest average scores across five metrics, including EN, SD, CC, SCD, and MS_SSIM. It also yields comparable results to the state-of-the-art methods in VIF and QAB/F, which verifies its capability in pixel intensity preservation, texture detail retention, and edge optimization. Scene 4-Scene 6 in Fig.8 display the qualitative evaluation results on the M3FD dataset, indicating that the method outperforms seven comparative methods in terms of artifact suppression and feature preservation. Quantitative results in Tab.2 show that the method obtains the highest mean values in CC, VIF, and MS_SSIM, while its performance in SCD and QAB/F is close to the optimal level of comparative methods. Overall, the method realizes the collaborative optimization of thermal target saliency and texture details via an adaptive fusion mechanism, thus verifying its technical effectiveness from multiple dimensions. Conclusions We propose a multimodal image fusion method based on SAMFuse. The fused images simultaneously retain the salient targets and detailed information of both modalities, delivering distinct advantages in overall visual quality. Quantitative results show that, compared with the optimal results of comparative algorithms, our method achieves respective improvements of 2.2%, 3.1% and 4.0% in SD, SCD and MS_SSIM on the TNO dataset; and increases of 7.9%, 4.6% and 3.1% in MS_SSIM, CC and VIF on the M3FD dataset. Furthermore, ablation experiments are conducted to verify the effectiveness of each core module proposed in this study.

Translated title of the contributionMultimodal image fusion method based on second-order attention and mixed-scale feature compensation
Original languageChinese (Traditional)
Article number20250514
JournalHongwai yu Jiguang Gongcheng/Infrared and Laser Engineering
Volume55
Issue number6
DOIs
Publication statusPublished - 25 Jun 2026

Fingerprint

Dive into the research topics of 'Multimodal image fusion method based on second-order attention and mixed-scale feature compensation'. Together they form a unique fingerprint.

Cite this