Abstract
Context-violating images contain visual information that is consistent with common sense but conflicts with a given context. For example, given the story of “Little Red Riding Hood”, an image depicting an old lady in a red hat finding a squirrel in the woods is a context-violating image, even though it looks visually plausible without the background story. Humans can easily identify and explain whether an image is consistent with the implicit constraints of a specific context, but can Multimodal Large Language Models (MLLMs) achieve similar performance? To explore the contextual reasoning capacity of MLLMs, we construct ContextualBench, a benchmark dataset consisting of context-violating images generated by text-to-image models. Each image is associated with a specific context and several constraints, and six types of context are defined in total. We use 15 MLLMs for four reasoning tasks on ContextualBench, and the results demonstrate that they fail to accurately identify and explain context-violating images, significantly falling behind human performance. As a pioneering step toward enhancing the contextual reasoning capacity of MLLMs, we propose a framework that retrieves context-related knowledge from external resources and integrates it into the inference phase of MLLMs. Extensive experiments demonstrate the effectiveness of the proposed framework.
| Original language | English |
|---|---|
| Article number | 114023 |
| Journal | Pattern Recognition |
| Volume | 180 |
| DOIs | |
| Publication status | Published - Dec 2026 |
| Externally published | Yes |
Keywords
- Context-violating images
- Contextual reasoning
- Multimodal large language models
Fingerprint
Dive into the research topics of 'Challenging and enhancing the reasoning capacity of multimodal LLMs in context-violating images'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver