TY - GEN
T1 - MSTMNet
T2 - 2025 China Automation Congress, CAC 2025
AU - Feng, Kaijing
AU - Li, Weixing
AU - Gao, Yan
AU - Zheng, Ziyi
AU - Wang, Ronghao
AU - Pan, Feng
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - In dynamic visual scenes, salient object detection faces critical challenges in maintaining temporal consistency and spatial accuracy, particularly manifested in insufficient spatiotemporal feature integration and inadequate modeling of long-term temporal dependencies. To tackle these challenges, we introduce the Multi-Scale Temporal Memory Network (MSTMNet), a novel framework that synergizes hierarchical spatiotemporal modeling with adaptive memory mechanisms. MSTMNet employs a dual-encoder architecture to separately extract features from the current query frame and historical memory frames, ensuring rich contextual representation. A dynamic memory bank with an attention-guided elimination policy selectively retains historically significant features based on relevance and freshness, while a Multi-Temporal Feature Fusion (MTFF) module aggregates multi-scale features across sequential temporal steps to capture both short-term motion cues and long-range dependencies. Subsequently, within the Spatial Feature Fusion (SFF) module, an enhanced self-attention mechanism refines low-level features by propagating semantic consistency from high-level representations. By combining memory-enhanced features with query features through a hierarchical decoder, MSTMNet achieves precise saliency maps. Comprehensive experiments on public datasets validate that MSTMNet surpasses transformer-based methods and outperforms established semi-supervised and optical-flow-dependent techniques in accuracy.
AB - In dynamic visual scenes, salient object detection faces critical challenges in maintaining temporal consistency and spatial accuracy, particularly manifested in insufficient spatiotemporal feature integration and inadequate modeling of long-term temporal dependencies. To tackle these challenges, we introduce the Multi-Scale Temporal Memory Network (MSTMNet), a novel framework that synergizes hierarchical spatiotemporal modeling with adaptive memory mechanisms. MSTMNet employs a dual-encoder architecture to separately extract features from the current query frame and historical memory frames, ensuring rich contextual representation. A dynamic memory bank with an attention-guided elimination policy selectively retains historically significant features based on relevance and freshness, while a Multi-Temporal Feature Fusion (MTFF) module aggregates multi-scale features across sequential temporal steps to capture both short-term motion cues and long-range dependencies. Subsequently, within the Spatial Feature Fusion (SFF) module, an enhanced self-attention mechanism refines low-level features by propagating semantic consistency from high-level representations. By combining memory-enhanced features with query features through a hierarchical decoder, MSTMNet achieves precise saliency maps. Comprehensive experiments on public datasets validate that MSTMNet surpasses transformer-based methods and outperforms established semi-supervised and optical-flow-dependent techniques in accuracy.
KW - dynamic memory bank
KW - multi-temporal feature fusion
KW - self-attention mechanism
KW - video salient object detection
UR - https://www.scopus.com/pages/publications/105041021084
U2 - 10.1109/CAC67268.2025.11487166
DO - 10.1109/CAC67268.2025.11487166
M3 - Conference contribution
AN - SCOPUS:105041021084
T3 - Proceedings - 2025 China Automation Congress, CAC 2025
SP - 5056
EP - 5061
BT - Proceedings - 2025 China Automation Congress, CAC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 26 September 2025 through 28 September 2025
ER -