Skip to main navigation Skip to search Skip to main content

Efficient Multi-agent Collaboration Learning via Posterior Mamba

  • Zhaohan Feng
  • , Wei Xiao
  • , Lei Yuan
  • , Yanjie Dong
  • , Gang Wang*
  • , Victor C.M. Leung
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • Nanjing University
  • Shenzhen MSU-BIT University
  • University of British Columbia

Research output: Contribution to journalArticlepeer-review

Abstract

Multi-agent systems increasingly arise in Big Data scenarios like distributed sensing, intelligent logistics, and multirobot data acquisition, where massive and heterogeneous data streams must be processed to support reliable collaborations. In such settings, agents must act based solely on local observations, leading to severe partial observability and non-stationarity that require modeling long-range temporal dependencies. Existing multi-agent reinforcement learning (MARL) approaches typically rely on recurrent neural networks (RNNs) or Transformers, yet remain limited by either representational capacity or computational efficiency. To address these challenges, we propose Posterior-Mamba (P-Mamba), a novel scalable recurrent MARL architecture that builds upon the recent Mamba state-space model and incorporates a posterior sampling mechanism to enhance robustness under stochastic and heterogeneous dynamics. P-Mamba supports efficient step-wise inference through a refactored recurrent execution path with latent-state caching, and can be extended by stacking multiple blocks. Extensive experiments on MPE, SMAC, SMACv2, and Multi-Agent Mu- JoCo show that P-Mamba consistently outperforms RNN- and Transformer-based policies. These results highlight P-Mamba as a strong candidate for data-intensive multi-agent systems requiring efficient, distributed, and long-horizon sequential decisionmaking. We also release the code of P-Mamba to facilitate reproducibility and broader use in the Big Data community at https://github.com/BUPT-zeld151/Posterior-Mamba.

Original languageEnglish
JournalIEEE Transactions on Big Data
DOIs
Publication statusAccepted/In press - 2026
Externally publishedYes

Keywords

  • Deep reinforcement learning
  • multi-agent coordination
  • non-stationary environments
  • posterior sampling
  • sequential decision-making
  • structured state space models

Fingerprint

Dive into the research topics of 'Efficient Multi-agent Collaboration Learning via Posterior Mamba'. Together they form a unique fingerprint.

Cite this