Wasserstein coupled graph learning for cross-modal retrieval

Yun Wang, Tong Zhang, Xueya Zhang, Zhen Cui*, Yuge Huang, Pengcheng Shen, Shaoxin Li, Jian Yang

*此作品的通讯作者

科研成果: 书/报告/会议事项章节会议稿件同行评审

19 引用 (Scopus)

摘要

Graphs play an important role in cross-modal image-text understanding as they characterize the intrinsic structure which is robust and crucial for the measurement of cross-modal similarity. In this work, we propose a Wasserstein Coupled Graph Learning (WCGL) method to deal with the cross-modal retrieval task. First, graphs are constructed according to two input cross-modal samples separately, and passed through the corresponding graph encoders to extract robust features. Then, a Wasserstein coupled dictionary, containing multiple pairs of counterpart graph keys with each key corresponding to one modality, is constructed for further feature learning. Based on this dictionary, the input graphs can be transformed into the dictionary space to facilitate the similarity measurement through a Wasserstein Graph Embedding (WGE) process. The WGE could capture the graph correlation between the input and each corresponding key through optimal transport, and hence well characterize the inter-graph structural relationship. To further achieve discriminant graph learning, we specifically define a Wasserstein discriminant loss on the coupled graph keys to make the intra-class (counterpart) keys more compact and inter-class (non-counterpart) keys more dispersed, which further promotes the final cross-modal retrieval task. Experimental results demonstrate the effectiveness and state-of-the-art performance.

源语言英语
主期刊名Proceedings - 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021
出版商Institute of Electrical and Electronics Engineers Inc.
1793-1802
页数10
ISBN(电子版)9781665428125
DOI
出版状态已出版 - 2021
已对外发布
活动18th IEEE/CVF International Conference on Computer Vision, ICCV 2021 - Virtual, Online, 加拿大
期限: 11 10月 202117 10月 2021

出版系列

姓名Proceedings of the IEEE International Conference on Computer Vision
ISSN(印刷版)1550-5499

会议

会议18th IEEE/CVF International Conference on Computer Vision, ICCV 2021
国家/地区加拿大
Virtual, Online
时期11/10/2117/10/21

指纹

探究 'Wasserstein coupled graph learning for cross-modal retrieval' 的科研主题。它们共同构成独一无二的指纹。

引用此