跳到主要导航 跳到搜索 跳到主要内容

An Efficient Multi-Agent Policy Self-Play Learning Method Aiming at Seize-Control Scenarios

  • Huaqing Zhang
  • , Hongbin Ma*
  • , Xiaofei Zhang
  • , Li Wang
  • , Minglei Han
  • , Hui Chen
  • , Ao Ding
  • *此作品的通讯作者
  • Beijing Institute of Technology
  • Tsinghua University

科研成果: 期刊稿件文章同行评审

摘要

Aiming at the problem of multi-Agent cooperative confrontation in seize-control scenarios, we design an efficient multi-Agent policy self-play (EMAP-SP) learning method. First, a multi-Agent centralized policy model is constructed to command the agents to perform tasks cooperatively. Considering that the policy being trained and its historical policies usually have poor exploration capability under incomplete information in self-play trainings, the intrinsic reward mechanism based on random network distillation (RND) is introduced in the self-play learning method. In addition, we propose a multi-step on-policy deep reinforcement learning (DRL) algorithm assisted by off-policy policy evaluation (MSOAO) to learn the best response policy in the self-play. Compared with DRL algorithms commonly used in complex decision problems, MSOAO has more efficient policy evaluation capability, and efficient policy evaluation further improves the policy learning capability. The effectiveness of EMAP-SP is fully verified in MiaoSuan wargame simulation system, and the evaluation results show that EMAP-SP can learn the cooperative policy of effectively defeating the Blue side's knowledge-based policy under incomplete information. Moreover, the evaluations results in DRL benchmark environments also show that the best response policy learning algorithm MSOAO can promote the agent to learn approximately optimal policies.

源语言英语
页(从-至)987-1004
页数18
期刊Unmanned Systems
13
4
DOI
出版状态已出版 - 1 7月 2025
已对外发布

学术指纹

探究 'An Efficient Multi-Agent Policy Self-Play Learning Method Aiming at Seize-Control Scenarios' 的科研主题。它们共同构成独一无二的学术指纹。

引用此