Abstract
Environmental non-stationarity remains a fundamental challenge in multi-agent reinforcement learning (MARL), hindering the efficiency of policy learning. Existing approaches primarily mitigate this issue by incorporating historical information into decisionmaking. However, widely adopted sequence modeling architectures, such as recurrent neural networks (RNNs) and Transformers, exhibit inherent limitations: RNNs struggle to capture long-range temporal dependencies, while Transformers, despite their superior encoding capabilities, suffer from the inflexibility imposed by a fixed-size context window and the quadratic computational complexity of self-attention. Motivated by the belief that self-supervised feature learning can enhance reinforcement learning (RL) efficiency, we propose self-predictive Mamba (SPMamba), a novel architecture that integrates Mamba’s superior sequence reasoning capabilities with a self-supervised auxiliary learning objective to facilitate the optimization of decentralized individual policies. Multiple challenging evaluations demonstrate that SPMamba is significantly superior to several state-of-the-art baselines.
| Original language | English |
|---|---|
| Article number | 192202 |
| Journal | Science China Information Sciences |
| Volume | 69 |
| Issue number | 9 |
| DOIs | |
| Publication status | Published - Sept 2026 |
Keywords
- multi-agent collaboration
- multi-agent reinforcement learning
- self-supervised learning
- structured state space model
Fingerprint
Dive into the research topics of 'Self-predictive Mamba for efficient multi-agent policy learning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver