跳到主要导航 跳到搜索 跳到主要内容

RGB-T Multi-modal Visual Question Answering in Nighttime and Adverse Environment

  • Songyuan Yang
  • , Fan Yang
  • , Biwen Yang
  • , Jing Zhao
  • , Yongqiang Sun*
  • , Ruiheng Zhang
  • *此作品的通讯作者
  • Beijing Institute of Technology
  • IFLYTEK Co., Ltd.
  • China Aerospace Science and Industry Corporation
  • Waterborne Transport Research Institute

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Visual Question Answering (VQA) models that rely only on RGB inputs often fail in nighttime and adverse environments due to poor illumination and semantic loss. To address this, we propose an RGB-T VQA framework that integrates visible and thermal infrared (TIR) modalities. The framework contains two key modules: a Cross-Modal Guided Attention (CGA) that uses thermal cues to refine RGB features, and a Thermal-Semantic Prior (TSP) that compensates for the limited semantics of TIR data. In addition, we construct a large-scale RGB-T VQA dataset covering diverse nighttime, low-light, and adverse weather scenes.

源语言英语
主期刊名Proceeding of the 2025 4th International Conference on Advanced Sensing and Intelligent Manufacturing, ASIM 2025
出版商Institute of Electrical and Electronics Engineers Inc.
ISBN(电子版)9798331554989
DOI
出版状态已出版 - 2025
活动4th International Conference on Advanced Sensing and Intelligent Manufacturing, ASIM 2025 - Changzhou, 中国
期限: 31 10月 20252 11月 2025

出版系列

姓名Proceeding of the 2025 4th International Conference on Advanced Sensing and Intelligent Manufacturing, ASIM 2025

会议

会议4th International Conference on Advanced Sensing and Intelligent Manufacturing, ASIM 2025
国家/地区中国
Changzhou
时期31/10/252/11/25

指纹

探究 'RGB-T Multi-modal Visual Question Answering in Nighttime and Adverse Environment' 的科研主题。它们共同构成独一无二的指纹。

引用此