Skip to main navigation Skip to search Skip to main content

RGB-T Multi-modal Visual Question Answering in Nighttime and Adverse Environment

  • Songyuan Yang
  • , Fan Yang
  • , Biwen Yang
  • , Jing Zhao
  • , Yongqiang Sun*
  • , Ruiheng Zhang
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • IFLYTEK Co., Ltd.
  • China Aerospace Science and Industry Corporation
  • Waterborne Transport Research Institute

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Visual Question Answering (VQA) models that rely only on RGB inputs often fail in nighttime and adverse environments due to poor illumination and semantic loss. To address this, we propose an RGB-T VQA framework that integrates visible and thermal infrared (TIR) modalities. The framework contains two key modules: a Cross-Modal Guided Attention (CGA) that uses thermal cues to refine RGB features, and a Thermal-Semantic Prior (TSP) that compensates for the limited semantics of TIR data. In addition, we construct a large-scale RGB-T VQA dataset covering diverse nighttime, low-light, and adverse weather scenes.

Original languageEnglish
Title of host publicationProceeding of the 2025 4th International Conference on Advanced Sensing and Intelligent Manufacturing, ASIM 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331554989
DOIs
Publication statusPublished - 2025
Event4th International Conference on Advanced Sensing and Intelligent Manufacturing, ASIM 2025 - Changzhou, China
Duration: 31 Oct 20252 Nov 2025

Publication series

NameProceeding of the 2025 4th International Conference on Advanced Sensing and Intelligent Manufacturing, ASIM 2025

Conference

Conference4th International Conference on Advanced Sensing and Intelligent Manufacturing, ASIM 2025
Country/TerritoryChina
CityChangzhou
Period31/10/252/11/25

Keywords

  • RGB-T fusion
  • Visual question answering (VQA)
  • low-light vision
  • multimodal learning
  • thermal infrared (TIR)

Fingerprint

Dive into the research topics of 'RGB-T Multi-modal Visual Question Answering in Nighttime and Adverse Environment'. Together they form a unique fingerprint.

Cite this