TY - GEN
T1 - Adaptive Knowledge Generation via Reinforcement-Guided Pattern Completion for Zero-Shot Visual Question Answering
AU - Sun, Zhihui
AU - Su, Diwei
AU - Li, Xiuxing
AU - Wang, Qixin
AU - Zhang, Shihao
AU - Wu, Xia
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/6/15
Y1 - 2026/6/15
N2 - Zero-shot visual question answering (VQA) requires models to reason over unseen image-question pairs without task-specific supervision, where performance is often undermined by incomplete visual grounding and unstable reasoning. Recent approaches attempt to mitigate this by using large language models to generate external knowledge conditioned on captions and questions. However, these methods typically operate in an open-loop manner, lacking mechanisms to regulate competing interpretations or correct misaligned knowledge. Inspired by the hippocampal pattern completion mechanism, which supports inference from partial observations through memory reactivation and competitive stabilization, we propose ARK-PC for zero-shot VQA, which formulates reasoning as a two-stage completion process. It first expands fragmented multimodal cues into multiple structured candidate hypotheses, then adaptively reinforces coherent candidates while suppressing inconsistent ones through iterative feedback. By coupling knowledge generation with competitive refinement, ARK-PC transforms open-loop inference into a closed-loop stabilization process, enabling robust reasoning under uncertainty without external supervision. Experiments on OK-VQA and A-OKVQA demonstrate consistent state-of-the-art zero-shot performance and strong generalization across diverse backbones, indicating that the improvements stem from the proposed framework rather than model scale. Figure 1:Comparison of typical zero-shot VQA paradigms. (a) Representation transfer methods adopt pretrained vision-language embeddings. (b) One-shot knowledge-augmented methods use LLMs/MLLMs to generate external knowledge in open-loop. (c) Our closed-loop cue-driven knowledge completion enables dynamic selection and stabilization of context-matched knowledge.
AB - Zero-shot visual question answering (VQA) requires models to reason over unseen image-question pairs without task-specific supervision, where performance is often undermined by incomplete visual grounding and unstable reasoning. Recent approaches attempt to mitigate this by using large language models to generate external knowledge conditioned on captions and questions. However, these methods typically operate in an open-loop manner, lacking mechanisms to regulate competing interpretations or correct misaligned knowledge. Inspired by the hippocampal pattern completion mechanism, which supports inference from partial observations through memory reactivation and competitive stabilization, we propose ARK-PC for zero-shot VQA, which formulates reasoning as a two-stage completion process. It first expands fragmented multimodal cues into multiple structured candidate hypotheses, then adaptively reinforces coherent candidates while suppressing inconsistent ones through iterative feedback. By coupling knowledge generation with competitive refinement, ARK-PC transforms open-loop inference into a closed-loop stabilization process, enabling robust reasoning under uncertainty without external supervision. Experiments on OK-VQA and A-OKVQA demonstrate consistent state-of-the-art zero-shot performance and strong generalization across diverse backbones, indicating that the improvements stem from the proposed framework rather than model scale. Figure 1:Comparison of typical zero-shot VQA paradigms. (a) Representation transfer methods adopt pretrained vision-language embeddings. (b) One-shot knowledge-augmented methods use LLMs/MLLMs to generate external knowledge in open-loop. (c) Our closed-loop cue-driven knowledge completion enables dynamic selection and stabilization of context-matched knowledge.
KW - Brain-inspired Computing
KW - Knowledge-Augmented Reasoning
KW - Visual Question Answering
KW - Zero-shot Learning
UR - https://www.scopus.com/pages/publications/105043401384
U2 - 10.1145/3805622.3810578
DO - 10.1145/3805622.3810578
M3 - Conference contribution
AN - SCOPUS:105043401384
T3 - ICMR 2026 - Proceedings of the 16th ACM International Conference on Multimedia Retrieval
SP - 1174
EP - 1183
BT - ICMR 2026 - Proceedings of the 16th ACM International Conference on Multimedia Retrieval
PB - Association for Computing Machinery, Inc
T2 - 16th ACM International Conference on Multimedia Retrieval, ICMR 2026
Y2 - 16 June 2026 through 19 June 2026
ER -