Skip to main navigation Skip to search Skip to main content

Self-predictive Mamba for efficient multi-agent policy learning

  • Beijing Institute of Technology
  • Beijing Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Environmental non-stationarity remains a fundamental challenge in multi-agent reinforcement learning (MARL), hindering the efficiency of policy learning. Existing approaches primarily mitigate this issue by incorporating historical information into decisionmaking. However, widely adopted sequence modeling architectures, such as recurrent neural networks (RNNs) and Transformers, exhibit inherent limitations: RNNs struggle to capture long-range temporal dependencies, while Transformers, despite their superior encoding capabilities, suffer from the inflexibility imposed by a fixed-size context window and the quadratic computational complexity of self-attention. Motivated by the belief that self-supervised feature learning can enhance reinforcement learning (RL) efficiency, we propose self-predictive Mamba (SPMamba), a novel architecture that integrates Mamba’s superior sequence reasoning capabilities with a self-supervised auxiliary learning objective to facilitate the optimization of decentralized individual policies. Multiple challenging evaluations demonstrate that SPMamba is significantly superior to several state-of-the-art baselines.

Original languageEnglish
Article number192202
JournalScience China Information Sciences
Volume69
Issue number9
DOIs
Publication statusPublished - Sept 2026

Keywords

  • multi-agent collaboration
  • multi-agent reinforcement learning
  • self-supervised learning
  • structured state space model

Fingerprint

Dive into the research topics of 'Self-predictive Mamba for efficient multi-agent policy learning'. Together they form a unique fingerprint.

Cite this