TY - GEN
T1 - A Hardware Accelerator for Infrared and Visible Image Fusion Based on Deep Learning
AU - Li, Rui
AU - Xie, Min
AU - Ma, Zhifeng
AU - Sun, Qian
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Most deep learning-based infrared and visible image fusion methods prioritize fusion quality while neglecting real-time requirements. Although lightweight CNN-based models reduce complexity, their limited ability to capture global dependency often leads to suboptimal fusion performance. This paper builds upon the lightweight image fusion network APWNet and incorporates Swin Transformer to enhance global modeling capability and improve fusion quality. Meanwhile, the computational bottlenecks in Swin Transformer are addressed through hardware acceleration. First, memory access bandwidth pressure in the shifted window attention module is alleviated via optimized data loading and an efficient masking mechanism. Second, computational parallelism in multi-head attention is improved by applying array partitioning to matrix multiplication. Third, a hardware-friendly pipelined softmax is designed to reduce resource consumption while maintaining computational accuracy. Our design is evaluated on the Xilinx xczu15eg FPGA. For infrared and visible images with a resolution of 256×256, the system achieves a processing time of approximately 28.5 ms. Although slightly slower than a GPU implementation, it provides a 2.4× improvement in energy efficiency.
AB - Most deep learning-based infrared and visible image fusion methods prioritize fusion quality while neglecting real-time requirements. Although lightweight CNN-based models reduce complexity, their limited ability to capture global dependency often leads to suboptimal fusion performance. This paper builds upon the lightweight image fusion network APWNet and incorporates Swin Transformer to enhance global modeling capability and improve fusion quality. Meanwhile, the computational bottlenecks in Swin Transformer are addressed through hardware acceleration. First, memory access bandwidth pressure in the shifted window attention module is alleviated via optimized data loading and an efficient masking mechanism. Second, computational parallelism in multi-head attention is improved by applying array partitioning to matrix multiplication. Third, a hardware-friendly pipelined softmax is designed to reduce resource consumption while maintaining computational accuracy. Our design is evaluated on the Xilinx xczu15eg FPGA. For infrared and visible images with a resolution of 256×256, the system achieves a processing time of approximately 28.5 ms. Although slightly slower than a GPU implementation, it provides a 2.4× improvement in energy efficiency.
KW - FPGA
KW - Hardware Accelerator
KW - Infrared and Visible Image Fusion
KW - Lightweight
UR - https://www.scopus.com/pages/publications/105043767463
U2 - 10.1109/AINIT70033.2026.11557798
DO - 10.1109/AINIT70033.2026.11557798
M3 - Conference contribution
AN - SCOPUS:105043767463
T3 - 2026 7th International Seminar on Artificial Intelligence, Networking and Information Technology, AINIT 2026
SP - 729
EP - 733
BT - 2026 7th International Seminar on Artificial Intelligence, Networking and Information Technology, AINIT 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 7th International Seminar on Artificial Intelligence, Networking and Information Technology, AINIT 2026
Y2 - 15 May 2026 through 17 May 2026
ER -