跳到主要导航 跳到搜索 跳到主要内容

Challenging and enhancing the reasoning capacity of multimodal LLMs in context-violating images

  • Yuyang Chen
  • , Hongxi Li
  • , Qiyuan Cheng
  • , Xinxiao Wu*
  • *此作品的通讯作者
  • Beijing Institute of Technology
  • Shenzhen MSU-BIT University

科研成果: 期刊稿件文章同行评审

摘要

Context-violating images contain visual information that is consistent with common sense but conflicts with a given context. For example, given the story of “Little Red Riding Hood”, an image depicting an old lady in a red hat finding a squirrel in the woods is a context-violating image, even though it looks visually plausible without the background story. Humans can easily identify and explain whether an image is consistent with the implicit constraints of a specific context, but can Multimodal Large Language Models (MLLMs) achieve similar performance? To explore the contextual reasoning capacity of MLLMs, we construct ContextualBench, a benchmark dataset consisting of context-violating images generated by text-to-image models. Each image is associated with a specific context and several constraints, and six types of context are defined in total. We use 15 MLLMs for four reasoning tasks on ContextualBench, and the results demonstrate that they fail to accurately identify and explain context-violating images, significantly falling behind human performance. As a pioneering step toward enhancing the contextual reasoning capacity of MLLMs, we propose a framework that retrieves context-related knowledge from external resources and integrates it into the inference phase of MLLMs. Extensive experiments demonstrate the effectiveness of the proposed framework.

源语言英语
文章编号114023
期刊Pattern Recognition
180
DOI
出版状态已出版 - 12月 2026
已对外发布

指纹

探究 'Challenging and enhancing the reasoning capacity of multimodal LLMs in context-violating images' 的科研主题。它们共同构成独一无二的指纹。

引用此