TY - GEN
T1 - Aggregation Enhanced Momentum Contrastive Learning for Unsupervised Cross-Modal Hashing
AU - Lu, Bo
AU - Zhao, Tianbao
AU - Liang, Guiyuan
AU - Ding, Xueyan
AU - Wang, Cunrui
AU - Yuan, Ye
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - Unsupervised cross-modal hashing has garnered widespread attention for its support of large-scale cross-modal retrieval. However, the exploration and preservation of inter-modal and intra-modal semantic structures among multi-modal instances remain limited, while the efficiency of learning and the performance of retrieval are affected by modality imbalances and uneven data distributions. In this paper, we propose a novel unsupervised cross-modal hash learning framework, namely Aggregation Enhanced Momentum Contrastive Cross-Modal Hashing (AEMCCH). Firstly, A multi-modal adjacency graph structure is proposed to construct heterogeneous multi-modal semantics correlation, in which significantly enhances the capability of capture and utilization of both global and local information from multi-modal data. Additionally, we introduce momentum contrastive learning to address the issues of data diversity and distribution disparity. Specifically, the hash queue dictionary is constructed by aggregating momentum features obtained from momentum encoders with the designed multi-modal adjacency graph. Simultaneously, a task-specific momentum updating strategy is proposed to ensure encoding consistency among the keys within the hash queue. Sufficient experiments on three benchmark datasets demonstrate that the proposed AEMCCH outperforms existing advanced unsupervised cross-modal hashing methods.
AB - Unsupervised cross-modal hashing has garnered widespread attention for its support of large-scale cross-modal retrieval. However, the exploration and preservation of inter-modal and intra-modal semantic structures among multi-modal instances remain limited, while the efficiency of learning and the performance of retrieval are affected by modality imbalances and uneven data distributions. In this paper, we propose a novel unsupervised cross-modal hash learning framework, namely Aggregation Enhanced Momentum Contrastive Cross-Modal Hashing (AEMCCH). Firstly, A multi-modal adjacency graph structure is proposed to construct heterogeneous multi-modal semantics correlation, in which significantly enhances the capability of capture and utilization of both global and local information from multi-modal data. Additionally, we introduce momentum contrastive learning to address the issues of data diversity and distribution disparity. Specifically, the hash queue dictionary is constructed by aggregating momentum features obtained from momentum encoders with the designed multi-modal adjacency graph. Simultaneously, a task-specific momentum updating strategy is proposed to ensure encoding consistency among the keys within the hash queue. Sufficient experiments on three benchmark datasets demonstrate that the proposed AEMCCH outperforms existing advanced unsupervised cross-modal hashing methods.
KW - cross-modal retrieval
KW - graph convolutional networks
KW - momentum contrastive learning
KW - unsupervised cross-modal hashing
UR - https://www.scopus.com/pages/publications/105029630197
U2 - 10.1007/978-981-95-5640-3_26
DO - 10.1007/978-981-95-5640-3_26
M3 - Conference contribution
AN - SCOPUS:105029630197
SN - 9789819556397
T3 - Lecture Notes in Computer Science
SP - 403
EP - 418
BT - Web and Big Data - 9th International Joint Conference, APWeb-WAIM 2025, Proceedings
A2 - Li, Jiajia
A2 - Chbeir, Richard
A2 - Li, Lei
A2 - Zong, Chuanyu
A2 - Zhang, Yanfeng
A2 - Zhang, Mengxuan
PB - Springer Science and Business Media Deutschland GmbH
T2 - 9th Asia-Pacific Web and Web-Age Information Management Joint International Conference on Web and Big Data, APWeb-WAIM 2025
Y2 - 28 August 2025 through 30 August 2025
ER -