TY - JOUR
T1 - Self-predictive Mamba for efficient multi-agent policy learning
AU - Feng, Zhaohan
AU - Wang, Runqing
AU - Zhang, Boxuan
AU - Sun, Jian
AU - Deng, Fang
AU - Wang, Gang
N1 - Publisher Copyright:
© Science China Press 2026.
PY - 2026/9
Y1 - 2026/9
N2 - Environmental non-stationarity remains a fundamental challenge in multi-agent reinforcement learning (MARL), hindering the efficiency of policy learning. Existing approaches primarily mitigate this issue by incorporating historical information into decisionmaking. However, widely adopted sequence modeling architectures, such as recurrent neural networks (RNNs) and Transformers, exhibit inherent limitations: RNNs struggle to capture long-range temporal dependencies, while Transformers, despite their superior encoding capabilities, suffer from the inflexibility imposed by a fixed-size context window and the quadratic computational complexity of self-attention. Motivated by the belief that self-supervised feature learning can enhance reinforcement learning (RL) efficiency, we propose self-predictive Mamba (SPMamba), a novel architecture that integrates Mamba’s superior sequence reasoning capabilities with a self-supervised auxiliary learning objective to facilitate the optimization of decentralized individual policies. Multiple challenging evaluations demonstrate that SPMamba is significantly superior to several state-of-the-art baselines.
AB - Environmental non-stationarity remains a fundamental challenge in multi-agent reinforcement learning (MARL), hindering the efficiency of policy learning. Existing approaches primarily mitigate this issue by incorporating historical information into decisionmaking. However, widely adopted sequence modeling architectures, such as recurrent neural networks (RNNs) and Transformers, exhibit inherent limitations: RNNs struggle to capture long-range temporal dependencies, while Transformers, despite their superior encoding capabilities, suffer from the inflexibility imposed by a fixed-size context window and the quadratic computational complexity of self-attention. Motivated by the belief that self-supervised feature learning can enhance reinforcement learning (RL) efficiency, we propose self-predictive Mamba (SPMamba), a novel architecture that integrates Mamba’s superior sequence reasoning capabilities with a self-supervised auxiliary learning objective to facilitate the optimization of decentralized individual policies. Multiple challenging evaluations demonstrate that SPMamba is significantly superior to several state-of-the-art baselines.
KW - multi-agent collaboration
KW - multi-agent reinforcement learning
KW - self-supervised learning
KW - structured state space model
UR - https://www.scopus.com/pages/publications/105040681575
U2 - 10.1007/s11432-025-4850-y
DO - 10.1007/s11432-025-4850-y
M3 - Article
AN - SCOPUS:105040681575
SN - 1674-733X
VL - 69
JO - Science China Information Sciences
JF - Science China Information Sciences
IS - 9
M1 - 192202
ER -