Skip to main navigation Skip to search Skip to main content

IRGPT: Understanding Real-World Infrared Image with Bi-Cross-Modal Curriculum on Large-Scale Benchmark

  • Zhe Cao
  • , Jin Zhang
  • , Ruiheng Zhang*
  • *Corresponding author for this work
  • Beijing Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Real-world infrared imagery presents unique challenges for vision-language models due to the scarcity of aligned text data and domain-specific characteristics. Although existing methods have advanced the field, their reliance on synthetic infrared images generated through style transfer from visible images, which limits their ability to capture the unique characteristics of the infrared modality. To address this, we propose IRGPT, the first multi-modal large language model for real-world infrared images, built upon a large-scale InfraRed-Text Dataset (IR-TD) comprising over 260 K authentic image-text pairs. The proposed IR-TD dataset contains real infrared images paired with meticulously handcrafted texts, where the initial drafts originated from two complementary processes: (1) LLM-generated descriptions of visible images, and (2) rule-based descriptions of annotations. Furthermore, we introduce a bi-cross-modal curriculum transfer learning strategy that systematically transfers knowledge from visible to infrared domains by considering the difficulty scores of both infrared-visible and infraredtext. Evaluated on a benchmark of 9 tasks (e.g., recognition, grounding), IRGPT achieves state-of-the-art performance even compared with larger-scale models.

Original languageEnglish
Title of host publicationProceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages166-176
Number of pages11
ISBN (Electronic)9798331587758
DOIs
Publication statusPublished - 2025
Event2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025 - Honolulu, United States
Duration: 19 Oct 202523 Oct 2025

Publication series

NameProceedings of the IEEE International Conference on Computer Vision
ISSN (Print)1550-5499
ISSN (Electronic)2380-7504

Conference

Conference2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
Country/TerritoryUnited States
CityHonolulu
Period19/10/2523/10/25

Keywords

  • benchmark
  • curriculum learning
  • infrared
  • multimodal large language model
  • vision language model

Fingerprint

Dive into the research topics of 'IRGPT: Understanding Real-World Infrared Image with Bi-Cross-Modal Curriculum on Large-Scale Benchmark'. Together they form a unique fingerprint.

Cite this