TY - JOUR
T1 - Challenging and enhancing the reasoning capacity of multimodal LLMs in context-violating images
AU - Chen, Yuyang
AU - Li, Hongxi
AU - Cheng, Qiyuan
AU - Wu, Xinxiao
N1 - Publisher Copyright:
© 2026 Elsevier Ltd
PY - 2026/12
Y1 - 2026/12
N2 - Context-violating images contain visual information that is consistent with common sense but conflicts with a given context. For example, given the story of “Little Red Riding Hood”, an image depicting an old lady in a red hat finding a squirrel in the woods is a context-violating image, even though it looks visually plausible without the background story. Humans can easily identify and explain whether an image is consistent with the implicit constraints of a specific context, but can Multimodal Large Language Models (MLLMs) achieve similar performance? To explore the contextual reasoning capacity of MLLMs, we construct ContextualBench, a benchmark dataset consisting of context-violating images generated by text-to-image models. Each image is associated with a specific context and several constraints, and six types of context are defined in total. We use 15 MLLMs for four reasoning tasks on ContextualBench, and the results demonstrate that they fail to accurately identify and explain context-violating images, significantly falling behind human performance. As a pioneering step toward enhancing the contextual reasoning capacity of MLLMs, we propose a framework that retrieves context-related knowledge from external resources and integrates it into the inference phase of MLLMs. Extensive experiments demonstrate the effectiveness of the proposed framework.
AB - Context-violating images contain visual information that is consistent with common sense but conflicts with a given context. For example, given the story of “Little Red Riding Hood”, an image depicting an old lady in a red hat finding a squirrel in the woods is a context-violating image, even though it looks visually plausible without the background story. Humans can easily identify and explain whether an image is consistent with the implicit constraints of a specific context, but can Multimodal Large Language Models (MLLMs) achieve similar performance? To explore the contextual reasoning capacity of MLLMs, we construct ContextualBench, a benchmark dataset consisting of context-violating images generated by text-to-image models. Each image is associated with a specific context and several constraints, and six types of context are defined in total. We use 15 MLLMs for four reasoning tasks on ContextualBench, and the results demonstrate that they fail to accurately identify and explain context-violating images, significantly falling behind human performance. As a pioneering step toward enhancing the contextual reasoning capacity of MLLMs, we propose a framework that retrieves context-related knowledge from external resources and integrates it into the inference phase of MLLMs. Extensive experiments demonstrate the effectiveness of the proposed framework.
KW - Context-violating images
KW - Contextual reasoning
KW - Multimodal large language models
UR - https://www.scopus.com/pages/publications/105041207519
U2 - 10.1016/j.patcog.2026.114023
DO - 10.1016/j.patcog.2026.114023
M3 - Article
AN - SCOPUS:105041207519
SN - 0031-3203
VL - 180
JO - Pattern Recognition
JF - Pattern Recognition
M1 - 114023
ER -