跳到主要导航 跳到搜索 跳到主要内容

Multi-scale dual-stream visual feature extraction and graph reasoning for visual question answering

  • Abdulganiyu Abdu Yusuf
  • , Chong Feng*
  • , Xianling Mao
  • , Xinyan Li
  • , Yunusa Haruna
  • , Ramadhani Ally Duma
  • *此作品的通讯作者
  • Beijing Institute of Technology
  • National Biotechnology Development Agency, Nigeria
  • Beijing Engineering Research Center of High Volume Language Information Processing and Cloud Computing Applications
  • China North Vehicle Research Institute
  • Beihang University

科研成果: 期刊稿件文章同行评审

摘要

Recent advancements in deep learning algorithms have significantly expanded the capabilities of systems to handle vision-to-language (V2L) tasks. Visual question answering (VQA) presents challenges that require a deep understanding of visual and language content to perform complex reasoning tasks. The existing VQA models often rely on grid-based or region-based visual features, which capture global context and object-specific details, respectively. However, balancing the complementary strengths of each feature type while minimizing fusion noise remains a significant challenge. This study propose a multi-scale dual-stream visual feature extraction method that combines grid and region features to enhance both global and local visual feature representations. Also, a visual graph relational reasoning (VGRR) approach is proposed to further improve reasoning by constructing a graph that models spatial and semantic relationships between visual objects, using Graph Attention Networks (GATs) for relational reasoning. To enhance the interaction between visual and textual modalities, we further propose a cross-modal self-attention fusion strategy, which enables the model to focus selectively on the most relevant parts of both the image and the question. The proposed model is evaluated on the VQA 2.0 and GQA benchmark datasets, demonstrating competitive performance with significant accuracy improvements compared to state-of-the-art methods. Ablation studies confirm the effectiveness of each module in enhancing visual-textual understanding and answer prediction.

源语言英语
文章编号544
期刊Applied Intelligence
55
6
DOI
出版状态已出版 - 4月 2025

学术指纹

探究 'Multi-scale dual-stream visual feature extraction and graph reasoning for visual question answering' 的科研主题。它们共同构成独一无二的学术指纹。

引用此