Skip to main navigation Skip to search Skip to main content

Adaptive Knowledge Generation via Reinforcement-Guided Pattern Completion for Zero-Shot Visual Question Answering

  • Zhihui Sun
  • , Diwei Su
  • , Xiuxing Li*
  • , Qixin Wang
  • , Shihao Zhang
  • , Xia Wu
  • *Corresponding author for this work
  • Beijing Normal University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Zero-shot visual question answering (VQA) requires models to reason over unseen image-question pairs without task-specific supervision, where performance is often undermined by incomplete visual grounding and unstable reasoning. Recent approaches attempt to mitigate this by using large language models to generate external knowledge conditioned on captions and questions. However, these methods typically operate in an open-loop manner, lacking mechanisms to regulate competing interpretations or correct misaligned knowledge. Inspired by the hippocampal pattern completion mechanism, which supports inference from partial observations through memory reactivation and competitive stabilization, we propose ARK-PC for zero-shot VQA, which formulates reasoning as a two-stage completion process. It first expands fragmented multimodal cues into multiple structured candidate hypotheses, then adaptively reinforces coherent candidates while suppressing inconsistent ones through iterative feedback. By coupling knowledge generation with competitive refinement, ARK-PC transforms open-loop inference into a closed-loop stabilization process, enabling robust reasoning under uncertainty without external supervision. Experiments on OK-VQA and A-OKVQA demonstrate consistent state-of-the-art zero-shot performance and strong generalization across diverse backbones, indicating that the improvements stem from the proposed framework rather than model scale. Figure 1:Comparison of typical zero-shot VQA paradigms. (a) Representation transfer methods adopt pretrained vision-language embeddings. (b) One-shot knowledge-augmented methods use LLMs/MLLMs to generate external knowledge in open-loop. (c) Our closed-loop cue-driven knowledge completion enables dynamic selection and stabilization of context-matched knowledge.

Original languageEnglish
Title of host publicationICMR 2026 - Proceedings of the 16th ACM International Conference on Multimedia Retrieval
PublisherAssociation for Computing Machinery, Inc
Pages1174-1183
Number of pages10
ISBN (Electronic)9798400726170
DOIs
Publication statusPublished - 15 Jun 2026
Externally publishedYes
Event16th ACM International Conference on Multimedia Retrieval, ICMR 2026 - Hybrid, Amsterdam, Netherlands
Duration: 16 Jun 202619 Jun 2026

Publication series

NameICMR 2026 - Proceedings of the 16th ACM International Conference on Multimedia Retrieval

Conference

Conference16th ACM International Conference on Multimedia Retrieval, ICMR 2026
Country/TerritoryNetherlands
CityHybrid, Amsterdam
Period16/06/2619/06/26

Keywords

  • Brain-inspired Computing
  • Knowledge-Augmented Reasoning
  • Visual Question Answering
  • Zero-shot Learning

Fingerprint

Dive into the research topics of 'Adaptive Knowledge Generation via Reinforcement-Guided Pattern Completion for Zero-Shot Visual Question Answering'. Together they form a unique fingerprint.

Cite this