跳到主要导航 跳到搜索 跳到主要内容

Self-predictive Mamba for efficient multi-agent policy learning

  • Beijing Institute of Technology
  • Beijing Institute of Technology

科研成果: 期刊稿件文章同行评审

摘要

Environmental non-stationarity remains a fundamental challenge in multi-agent reinforcement learning (MARL), hindering the efficiency of policy learning. Existing approaches primarily mitigate this issue by incorporating historical information into decisionmaking. However, widely adopted sequence modeling architectures, such as recurrent neural networks (RNNs) and Transformers, exhibit inherent limitations: RNNs struggle to capture long-range temporal dependencies, while Transformers, despite their superior encoding capabilities, suffer from the inflexibility imposed by a fixed-size context window and the quadratic computational complexity of self-attention. Motivated by the belief that self-supervised feature learning can enhance reinforcement learning (RL) efficiency, we propose self-predictive Mamba (SPMamba), a novel architecture that integrates Mamba’s superior sequence reasoning capabilities with a self-supervised auxiliary learning objective to facilitate the optimization of decentralized individual policies. Multiple challenging evaluations demonstrate that SPMamba is significantly superior to several state-of-the-art baselines.

源语言英语
期刊论文编号192202
期刊Science China Information Sciences
69
9
DOI
出版状态已出版 - 9月 2026

学术指纹

探究 'Self-predictive Mamba for efficient multi-agent policy learning' 的科研主题。它们共同构成独一无二的学术指纹。

引用此