TY - GEN
T1 - IRGPT
T2 - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
AU - Cao, Zhe
AU - Zhang, Jin
AU - Zhang, Ruiheng
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Real-world infrared imagery presents unique challenges for vision-language models due to the scarcity of aligned text data and domain-specific characteristics. Although existing methods have advanced the field, their reliance on synthetic infrared images generated through style transfer from visible images, which limits their ability to capture the unique characteristics of the infrared modality. To address this, we propose IRGPT, the first multi-modal large language model for real-world infrared images, built upon a large-scale InfraRed-Text Dataset (IR-TD) comprising over 260 K authentic image-text pairs. The proposed IR-TD dataset contains real infrared images paired with meticulously handcrafted texts, where the initial drafts originated from two complementary processes: (1) LLM-generated descriptions of visible images, and (2) rule-based descriptions of annotations. Furthermore, we introduce a bi-cross-modal curriculum transfer learning strategy that systematically transfers knowledge from visible to infrared domains by considering the difficulty scores of both infrared-visible and infraredtext. Evaluated on a benchmark of 9 tasks (e.g., recognition, grounding), IRGPT achieves state-of-the-art performance even compared with larger-scale models.
AB - Real-world infrared imagery presents unique challenges for vision-language models due to the scarcity of aligned text data and domain-specific characteristics. Although existing methods have advanced the field, their reliance on synthetic infrared images generated through style transfer from visible images, which limits their ability to capture the unique characteristics of the infrared modality. To address this, we propose IRGPT, the first multi-modal large language model for real-world infrared images, built upon a large-scale InfraRed-Text Dataset (IR-TD) comprising over 260 K authentic image-text pairs. The proposed IR-TD dataset contains real infrared images paired with meticulously handcrafted texts, where the initial drafts originated from two complementary processes: (1) LLM-generated descriptions of visible images, and (2) rule-based descriptions of annotations. Furthermore, we introduce a bi-cross-modal curriculum transfer learning strategy that systematically transfers knowledge from visible to infrared domains by considering the difficulty scores of both infrared-visible and infraredtext. Evaluated on a benchmark of 9 tasks (e.g., recognition, grounding), IRGPT achieves state-of-the-art performance even compared with larger-scale models.
KW - benchmark
KW - curriculum learning
KW - infrared
KW - multimodal large language model
KW - vision language model
UR - https://www.scopus.com/pages/publications/105044246328
U2 - 10.1109/ICCV51701.2025.00023
DO - 10.1109/ICCV51701.2025.00023
M3 - Conference contribution
AN - SCOPUS:105044246328
T3 - Proceedings of the IEEE International Conference on Computer Vision
SP - 166
EP - 176
BT - Proceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 19 October 2025 through 23 October 2025
ER -