TY - JOUR
T1 - MPANet
T2 - Motion Pattern Aggregation Network for Gait Recognition
AU - Ma, Kang
AU - Liang, Xiao
AU - Cao, Chunshui
AU - Wang, Guoren
AU - Zheng, Dezhi
N1 - Publisher Copyright:
© The Author(s), under exclusive licence to Springer Science+Business Media, LLC, part of Springer Nature 2026.
PY - 2026/7
Y1 - 2026/7
N2 - Gait recognition is valuable in a variety of applications, including social security, crime investigation, and video surveillance. However, gait recognition faces numerous external factors in real-world scenarios, including wearing overcoats, carrying conditions, and diverse viewing angles. Furthermore, distractors such as occlusions, crowds, and directional changes further increase the complexity. In recent years, numerous deep learning-based gait recognition methods have shown promising performance. However, these methods often employ a convolutional network with fixed weights for feature extraction, which inadequately addresses the dynamic localization of key regions and the extraction of robust local motion patterns. Additionally, the aggregation of complete motion patterns is neglected. In this paper, we present a new perspective suggesting that gait features include global motion patterns across multiple key regions, with each global motion pattern composed of a series of local motion patterns. To this end, we introduce a Motion Pattern Aggregation Network aimed at learning more discriminative features. Specifically, we employ a dynamic attention mechanism among neighboring pixel features, which adaptively focuses on key regions and generates expressive local motion patterns. Furthermore, we introduce a novel self-attention mechanism to select representative local motion patterns and then aggregate them according to the Nyquist-Shannon sampling theorem to acquire global motion patterns. Extensive experiments conducted in both laboratory and real-world settings have verified the effectiveness of the proposed method and highlighted the necessity of investigating the intrinsic hierarchical structure of motion patterns.
AB - Gait recognition is valuable in a variety of applications, including social security, crime investigation, and video surveillance. However, gait recognition faces numerous external factors in real-world scenarios, including wearing overcoats, carrying conditions, and diverse viewing angles. Furthermore, distractors such as occlusions, crowds, and directional changes further increase the complexity. In recent years, numerous deep learning-based gait recognition methods have shown promising performance. However, these methods often employ a convolutional network with fixed weights for feature extraction, which inadequately addresses the dynamic localization of key regions and the extraction of robust local motion patterns. Additionally, the aggregation of complete motion patterns is neglected. In this paper, we present a new perspective suggesting that gait features include global motion patterns across multiple key regions, with each global motion pattern composed of a series of local motion patterns. To this end, we introduce a Motion Pattern Aggregation Network aimed at learning more discriminative features. Specifically, we employ a dynamic attention mechanism among neighboring pixel features, which adaptively focuses on key regions and generates expressive local motion patterns. Furthermore, we introduce a novel self-attention mechanism to select representative local motion patterns and then aggregate them according to the Nyquist-Shannon sampling theorem to acquire global motion patterns. Extensive experiments conducted in both laboratory and real-world settings have verified the effectiveness of the proposed method and highlighted the necessity of investigating the intrinsic hierarchical structure of motion patterns.
KW - Dynamic attention mechanism
KW - Gait recognition
KW - Motion pattern aggregation
KW - Self-attention mechanism
UR - https://www.scopus.com/pages/publications/105042392184
U2 - 10.1007/s11263-026-02912-1
DO - 10.1007/s11263-026-02912-1
M3 - Article
AN - SCOPUS:105042392184
SN - 0920-5691
VL - 134
JO - International Journal of Computer Vision
JF - International Journal of Computer Vision
IS - 7
M1 - 323
ER -