跳到主要导航 跳到搜索 跳到主要内容

Adaptive Knowledge Generation via Reinforcement-Guided Pattern Completion for Zero-Shot Visual Question Answering

  • Zhihui Sun
  • , Diwei Su
  • , Xiuxing Li*
  • , Qixin Wang
  • , Shihao Zhang
  • , Xia Wu
  • *此作品的通讯作者
  • Beijing Normal University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Zero-shot visual question answering (VQA) requires models to reason over unseen image-question pairs without task-specific supervision, where performance is often undermined by incomplete visual grounding and unstable reasoning. Recent approaches attempt to mitigate this by using large language models to generate external knowledge conditioned on captions and questions. However, these methods typically operate in an open-loop manner, lacking mechanisms to regulate competing interpretations or correct misaligned knowledge. Inspired by the hippocampal pattern completion mechanism, which supports inference from partial observations through memory reactivation and competitive stabilization, we propose ARK-PC for zero-shot VQA, which formulates reasoning as a two-stage completion process. It first expands fragmented multimodal cues into multiple structured candidate hypotheses, then adaptively reinforces coherent candidates while suppressing inconsistent ones through iterative feedback. By coupling knowledge generation with competitive refinement, ARK-PC transforms open-loop inference into a closed-loop stabilization process, enabling robust reasoning under uncertainty without external supervision. Experiments on OK-VQA and A-OKVQA demonstrate consistent state-of-the-art zero-shot performance and strong generalization across diverse backbones, indicating that the improvements stem from the proposed framework rather than model scale. Figure 1:Comparison of typical zero-shot VQA paradigms. (a) Representation transfer methods adopt pretrained vision-language embeddings. (b) One-shot knowledge-augmented methods use LLMs/MLLMs to generate external knowledge in open-loop. (c) Our closed-loop cue-driven knowledge completion enables dynamic selection and stabilization of context-matched knowledge.

源语言英语
主期刊名ICMR 2026 - Proceedings of the 16th ACM International Conference on Multimedia Retrieval
出版商Association for Computing Machinery, Inc
1174-1183
页数10
ISBN(电子版)9798400726170
DOI
出版状态已出版 - 15 6月 2026
已对外发布
活动16th ACM International Conference on Multimedia Retrieval, ICMR 2026 - Hybrid, Amsterdam, 荷兰
期限: 16 6月 202619 6月 2026

出版系列

姓名ICMR 2026 - Proceedings of the 16th ACM International Conference on Multimedia Retrieval

会议

会议16th ACM International Conference on Multimedia Retrieval, ICMR 2026
国家/地区荷兰
Hybrid, Amsterdam
时期16/06/2619/06/26

指纹

探究 'Adaptive Knowledge Generation via Reinforcement-Guided Pattern Completion for Zero-Shot Visual Question Answering' 的科研主题。它们共同构成独一无二的指纹。

引用此