Skip to main navigation Skip to search Skip to main content

MSTMNet: A Multi-Scale Temporal Memory Network with Dynamic Memory Bank for Video Salient Object Detection

  • Kaijing Feng
  • , Weixing Li
  • , Yan Gao
  • , Ziyi Zheng
  • , Ronghao Wang
  • , Feng Pan*
  • *Corresponding author for this work
  • Beijing Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In dynamic visual scenes, salient object detection faces critical challenges in maintaining temporal consistency and spatial accuracy, particularly manifested in insufficient spatiotemporal feature integration and inadequate modeling of long-term temporal dependencies. To tackle these challenges, we introduce the Multi-Scale Temporal Memory Network (MSTMNet), a novel framework that synergizes hierarchical spatiotemporal modeling with adaptive memory mechanisms. MSTMNet employs a dual-encoder architecture to separately extract features from the current query frame and historical memory frames, ensuring rich contextual representation. A dynamic memory bank with an attention-guided elimination policy selectively retains historically significant features based on relevance and freshness, while a Multi-Temporal Feature Fusion (MTFF) module aggregates multi-scale features across sequential temporal steps to capture both short-term motion cues and long-range dependencies. Subsequently, within the Spatial Feature Fusion (SFF) module, an enhanced self-attention mechanism refines low-level features by propagating semantic consistency from high-level representations. By combining memory-enhanced features with query features through a hierarchical decoder, MSTMNet achieves precise saliency maps. Comprehensive experiments on public datasets validate that MSTMNet surpasses transformer-based methods and outperforms established semi-supervised and optical-flow-dependent techniques in accuracy.

Original languageEnglish
Title of host publicationProceedings - 2025 China Automation Congress, CAC 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages5056-5061
Number of pages6
ISBN (Electronic)9798331589677
DOIs
Publication statusPublished - 2025
Externally publishedYes
Event2025 China Automation Congress, CAC 2025 - Harbin, China
Duration: 26 Sept 202528 Sept 2025

Publication series

NameProceedings - 2025 China Automation Congress, CAC 2025

Conference

Conference2025 China Automation Congress, CAC 2025
Country/TerritoryChina
CityHarbin
Period26/09/2528/09/25

Keywords

  • dynamic memory bank
  • multi-temporal feature fusion
  • self-attention mechanism
  • video salient object detection

Fingerprint

Dive into the research topics of 'MSTMNet: A Multi-Scale Temporal Memory Network with Dynamic Memory Bank for Video Salient Object Detection'. Together they form a unique fingerprint.

Cite this