Skip to main navigation Skip to search Skip to main content

Aggregation Enhanced Momentum Contrastive Learning for Unsupervised Cross-Modal Hashing

  • Bo Lu
  • , Tianbao Zhao*
  • , Guiyuan Liang
  • , Xueyan Ding
  • , Cunrui Wang
  • , Ye Yuan
  • *Corresponding author for this work
  • Dalian Minzu University
  • Beijing Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Unsupervised cross-modal hashing has garnered widespread attention for its support of large-scale cross-modal retrieval. However, the exploration and preservation of inter-modal and intra-modal semantic structures among multi-modal instances remain limited, while the efficiency of learning and the performance of retrieval are affected by modality imbalances and uneven data distributions. In this paper, we propose a novel unsupervised cross-modal hash learning framework, namely Aggregation Enhanced Momentum Contrastive Cross-Modal Hashing (AEMCCH). Firstly, A multi-modal adjacency graph structure is proposed to construct heterogeneous multi-modal semantics correlation, in which significantly enhances the capability of capture and utilization of both global and local information from multi-modal data. Additionally, we introduce momentum contrastive learning to address the issues of data diversity and distribution disparity. Specifically, the hash queue dictionary is constructed by aggregating momentum features obtained from momentum encoders with the designed multi-modal adjacency graph. Simultaneously, a task-specific momentum updating strategy is proposed to ensure encoding consistency among the keys within the hash queue. Sufficient experiments on three benchmark datasets demonstrate that the proposed AEMCCH outperforms existing advanced unsupervised cross-modal hashing methods.

Original languageEnglish
Title of host publicationWeb and Big Data - 9th International Joint Conference, APWeb-WAIM 2025, Proceedings
EditorsJiajia Li, Richard Chbeir, Lei Li, Chuanyu Zong, Yanfeng Zhang, Mengxuan Zhang
PublisherSpringer Science and Business Media Deutschland GmbH
Pages403-418
Number of pages16
ISBN (Print)9789819556397
DOIs
Publication statusPublished - 2026
Externally publishedYes
Event9th Asia-Pacific Web and Web-Age Information Management Joint International Conference on Web and Big Data, APWeb-WAIM 2025 - Shenyang, China
Duration: 28 Aug 202530 Aug 2025

Publication series

NameLecture Notes in Computer Science
Volume16113 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference9th Asia-Pacific Web and Web-Age Information Management Joint International Conference on Web and Big Data, APWeb-WAIM 2025
Country/TerritoryChina
CityShenyang
Period28/08/2530/08/25

Keywords

  • cross-modal retrieval
  • graph convolutional networks
  • momentum contrastive learning
  • unsupervised cross-modal hashing

Fingerprint

Dive into the research topics of 'Aggregation Enhanced Momentum Contrastive Learning for Unsupervised Cross-Modal Hashing'. Together they form a unique fingerprint.

Cite this