TY - GEN
T1 - Selective Pseudo Word Inversion with MLLM Reasoning for Zero-Shot Composed Image Retrieval
AU - Yu, Jing
AU - Ru, Zhipeng
AU - Gan, Minggang
AU - Yue, Zhao
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2027.
PY - 2027
Y1 - 2027
N2 - Composed Image Retrieval (CIR) enables users to search for target images via a multimodal query including a reference image and a modification text, aiming to retain key visual features while integrating textual modifications. Since supervised CIR requires costly annotated triplets, researchers have explored Zero-Shot CIR (ZS-CIR). Current ZS-CIR primarily includes two kinds of methods. Textual inversion methods map reference images to pseudo-word tokens, which often fail to understand manipulation intentions and retain excessive visual noise from the reference image. Training-free methods leverage Multimodal Large Language Models (MLLMs), which lose fine-grained visual details by expressing visual content via only text. To overcome these limitations, we propose a novel Selective Pseudo Word Inversion with MLLM Reasoning (SPIR), which leverages textual inversion to supplement visual details and enhances reasoning using MLLM-generated descriptions. To align the inference and training schema, we leverage MLLMs to process image-text pairs, automatically generating modification texts and corresponding target descriptions that serve as supervision information. Extensive experiments conducted on three benchmark datasets demonstrate the superiority of our proposed method. Our code is released at https://github.com/fancySummer19/SPIR.
AB - Composed Image Retrieval (CIR) enables users to search for target images via a multimodal query including a reference image and a modification text, aiming to retain key visual features while integrating textual modifications. Since supervised CIR requires costly annotated triplets, researchers have explored Zero-Shot CIR (ZS-CIR). Current ZS-CIR primarily includes two kinds of methods. Textual inversion methods map reference images to pseudo-word tokens, which often fail to understand manipulation intentions and retain excessive visual noise from the reference image. Training-free methods leverage Multimodal Large Language Models (MLLMs), which lose fine-grained visual details by expressing visual content via only text. To overcome these limitations, we propose a novel Selective Pseudo Word Inversion with MLLM Reasoning (SPIR), which leverages textual inversion to supplement visual details and enhances reasoning using MLLM-generated descriptions. To align the inference and training schema, we leverage MLLMs to process image-text pairs, automatically generating modification texts and corresponding target descriptions that serve as supervision information. Extensive experiments conducted on three benchmark datasets demonstrate the superiority of our proposed method. Our code is released at https://github.com/fancySummer19/SPIR.
KW - Composed image retrieval
KW - Contrastive Learning
KW - Multimodal large language models
UR - https://www.scopus.com/pages/publications/105046308715
U2 - 10.1007/978-981-92-2856-0_41
DO - 10.1007/978-981-92-2856-0_41
M3 - Conference contribution
AN - SCOPUS:105046308715
SN - 9789819228553
T3 - Lecture Notes in Computer Science
SP - 573
EP - 589
BT - Knowledge Science, Engineering and Management - 19th International Conference, KSEM 2026, Proceedings
A2 - Niu, Jianwei
A2 - Qiu, Meikang
A2 - Cao, Cungen
PB - Springer Science and Business Media Deutschland GmbH
T2 - 19th International Conference on Knowledge Science, Engineering and Management, KSEM 2026
Y2 - 17 July 2026 through 19 July 2026
ER -