VQ-InfraTrans: A Unified Framework for RGB-IR Translation with Hybrid Transformer

Qiyang Sun; Xia Wang; Changda Yan; Xin Zhang

doi:10.3390/rs15245661

VQ-InfraTrans: A Unified Framework for RGB-IR Translation with Hybrid Transformer

Qiyang Sun, Xia Wang^*, Changda Yan, Xin Zhang

^*此作品的通讯作者

光电学院

Beijing Institute of Technology

科研成果: 期刊稿件 › 文章 › 同行评审

摘要

Infrared (IR) images containing rich spectral information are essential in many fields. Most RGB-IR transfer work currently relies on conditional generative models to learn and train IR images for specific devices and scenes. However, these models only establish an empirical mapping relationship between RGB and IR images in a single dataset, which cannot achieve the multi-scene and multi-band (0.7–3 (Formula presented.) m and 8–15 (Formula presented.) m) transfer task. To address this challenge, we propose VQ-InfraTrans, a comprehensive framework for transferring images from the visible spectrum to the infrared spectrum. Our framework incorporates a multi-mode approach to RGB-IR image transferring, encompassing both unconditional and conditional transfers, achieving diverse and flexible image transformations. Instead of training individual models for each specific condition or dataset, we propose a two-stage transfer framework that integrates diverse requirements into a unified model that utilizes a composite encoder–decoder based on VQ-GAN, and a multi-path transformer to translate multi-modal images from RGB to infrared. To address the issue of significant errors in transferring specific targets due to their radiance, we have developed a hybrid editing module to precisely map spectral transfer information for specific local targets. The qualitative and quantitative comparisons conducted in this work reveal substantial enhancements compared to prior algorithms, as the objective evaluation metric SSIM (structural similarity index) was improved by 2.24% and the PSNR (peak signal-to-noise ratio) was improved by 2.71%.

源语言	英语
文章编号	5661
期刊	Remote Sensing
卷	15
期	24
DOI	https://doi.org/10.3390/rs15245661
出版状态	已出版 - 12月 2023

访问文件

10.3390/rs15245661

其它文件与链接

链接到 Scopus 的出版物

引用此

@article{2656c892357a4ae9be62539c89dca1ca,

title = "VQ-InfraTrans: A Unified Framework for RGB-IR Translation with Hybrid Transformer",

abstract = "Infrared (IR) images containing rich spectral information are essential in many fields. Most RGB-IR transfer work currently relies on conditional generative models to learn and train IR images for specific devices and scenes. However, these models only establish an empirical mapping relationship between RGB and IR images in a single dataset, which cannot achieve the multi-scene and multi-band (0.7–3 (Formula presented.) m and 8–15 (Formula presented.) m) transfer task. To address this challenge, we propose VQ-InfraTrans, a comprehensive framework for transferring images from the visible spectrum to the infrared spectrum. Our framework incorporates a multi-mode approach to RGB-IR image transferring, encompassing both unconditional and conditional transfers, achieving diverse and flexible image transformations. Instead of training individual models for each specific condition or dataset, we propose a two-stage transfer framework that integrates diverse requirements into a unified model that utilizes a composite encoder–decoder based on VQ-GAN, and a multi-path transformer to translate multi-modal images from RGB to infrared. To address the issue of significant errors in transferring specific targets due to their radiance, we have developed a hybrid editing module to precisely map spectral transfer information for specific local targets. The qualitative and quantitative comparisons conducted in this work reveal substantial enhancements compared to prior algorithms, as the objective evaluation metric SSIM (structural similarity index) was improved by 2.24% and the PSNR (peak signal-to-noise ratio) was improved by 2.71%.",

keywords = "image-to-image translation, infrared image, multi-modal controls, transformer, vector quantization",

author = "Qiyang Sun and Xia Wang and Changda Yan and Xin Zhang",

note = "Publisher Copyright: {\textcopyright} 2023 by the authors.",

year = "2023",

month = dec,

doi = "10.3390/rs15245661",

language = "English",

volume = "15",

journal = "Remote Sensing",

issn = "2072-4292",

publisher = "Multidisciplinary Digital Publishing Institute (MDPI)",

number = "24",

}

TY - JOUR

T1 - VQ-InfraTrans

T2 - A Unified Framework for RGB-IR Translation with Hybrid Transformer

AU - Sun, Qiyang

AU - Wang, Xia

AU - Yan, Changda

AU - Zhang, Xin

PY - 2023/12

Y1 - 2023/12

N2 - Infrared (IR) images containing rich spectral information are essential in many fields. Most RGB-IR transfer work currently relies on conditional generative models to learn and train IR images for specific devices and scenes. However, these models only establish an empirical mapping relationship between RGB and IR images in a single dataset, which cannot achieve the multi-scene and multi-band (0.7–3 (Formula presented.) m and 8–15 (Formula presented.) m) transfer task. To address this challenge, we propose VQ-InfraTrans, a comprehensive framework for transferring images from the visible spectrum to the infrared spectrum. Our framework incorporates a multi-mode approach to RGB-IR image transferring, encompassing both unconditional and conditional transfers, achieving diverse and flexible image transformations. Instead of training individual models for each specific condition or dataset, we propose a two-stage transfer framework that integrates diverse requirements into a unified model that utilizes a composite encoder–decoder based on VQ-GAN, and a multi-path transformer to translate multi-modal images from RGB to infrared. To address the issue of significant errors in transferring specific targets due to their radiance, we have developed a hybrid editing module to precisely map spectral transfer information for specific local targets. The qualitative and quantitative comparisons conducted in this work reveal substantial enhancements compared to prior algorithms, as the objective evaluation metric SSIM (structural similarity index) was improved by 2.24% and the PSNR (peak signal-to-noise ratio) was improved by 2.71%.

AB - Infrared (IR) images containing rich spectral information are essential in many fields. Most RGB-IR transfer work currently relies on conditional generative models to learn and train IR images for specific devices and scenes. However, these models only establish an empirical mapping relationship between RGB and IR images in a single dataset, which cannot achieve the multi-scene and multi-band (0.7–3 (Formula presented.) m and 8–15 (Formula presented.) m) transfer task. To address this challenge, we propose VQ-InfraTrans, a comprehensive framework for transferring images from the visible spectrum to the infrared spectrum. Our framework incorporates a multi-mode approach to RGB-IR image transferring, encompassing both unconditional and conditional transfers, achieving diverse and flexible image transformations. Instead of training individual models for each specific condition or dataset, we propose a two-stage transfer framework that integrates diverse requirements into a unified model that utilizes a composite encoder–decoder based on VQ-GAN, and a multi-path transformer to translate multi-modal images from RGB to infrared. To address the issue of significant errors in transferring specific targets due to their radiance, we have developed a hybrid editing module to precisely map spectral transfer information for specific local targets. The qualitative and quantitative comparisons conducted in this work reveal substantial enhancements compared to prior algorithms, as the objective evaluation metric SSIM (structural similarity index) was improved by 2.24% and the PSNR (peak signal-to-noise ratio) was improved by 2.71%.

KW - image-to-image translation

KW - infrared image

KW - multi-modal controls

KW - transformer

KW - vector quantization

UR - http://www.scopus.com/inward/record.url?scp=85180617028&partnerID=8YFLogxK

U2 - 10.3390/rs15245661

DO - 10.3390/rs15245661

M3 - Article

AN - SCOPUS:85180617028

SN - 2072-4292

VL - 15

JO - Remote Sensing

JF - Remote Sensing

IS - 24

M1 - 5661

ER -

VQ-InfraTrans: A Unified Framework for RGB-IR Translation with Hybrid Transformer

摘要

访问文件

其它文件与链接

指纹

引用此