Skip to main navigation Skip to search Skip to main content

Challenging and enhancing the reasoning capacity of multimodal LLMs in context-violating images

  • Yuyang Chen
  • , Hongxi Li
  • , Qiyuan Cheng
  • , Xinxiao Wu*
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • Shenzhen MSU-BIT University

Research output: Contribution to journalArticlepeer-review

Abstract

Context-violating images contain visual information that is consistent with common sense but conflicts with a given context. For example, given the story of “Little Red Riding Hood”, an image depicting an old lady in a red hat finding a squirrel in the woods is a context-violating image, even though it looks visually plausible without the background story. Humans can easily identify and explain whether an image is consistent with the implicit constraints of a specific context, but can Multimodal Large Language Models (MLLMs) achieve similar performance? To explore the contextual reasoning capacity of MLLMs, we construct ContextualBench, a benchmark dataset consisting of context-violating images generated by text-to-image models. Each image is associated with a specific context and several constraints, and six types of context are defined in total. We use 15 MLLMs for four reasoning tasks on ContextualBench, and the results demonstrate that they fail to accurately identify and explain context-violating images, significantly falling behind human performance. As a pioneering step toward enhancing the contextual reasoning capacity of MLLMs, we propose a framework that retrieves context-related knowledge from external resources and integrates it into the inference phase of MLLMs. Extensive experiments demonstrate the effectiveness of the proposed framework.

Original languageEnglish
Article number114023
JournalPattern Recognition
Volume180
DOIs
Publication statusPublished - Dec 2026
Externally publishedYes

Keywords

  • Context-violating images
  • Contextual reasoning
  • Multimodal large language models

Fingerprint

Dive into the research topics of 'Challenging and enhancing the reasoning capacity of multimodal LLMs in context-violating images'. Together they form a unique fingerprint.

Cite this