跳到主要导航 跳到搜索 跳到主要内容

Multi-time scale hierarchical trust domain leads to the improvement of MAPPO algorithm

  • Zhentao Guo
  • , Licheng Sun
  • , Guiyu Zhao
  • , Tianhao Wang
  • , Ao Ding
  • , Hongbin Ma*
  • *此作品的通讯作者
  • Beijing Institute of Technology

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Multi-agent Proximal Policy Optimization is a ubiquitous on-policy reinforcement learning algorithm, but its usage is significantly lower than that of off-policy learning algorithms in multi-agent environments. The existing MAPPO algorithm has the problem of insufficient generalization ability, adaptability and training stability when dealing with complex tasks. In this paper, we propose an improved trust domain guided MAPPO algorithm with multi-time scale hierarchical structure, which aims to cope with the dynamic changes of hierarchical structure and multi-time scale of tasks. A multi-time scale hierarchical structure is introduced by the algorithm, along with trust domain constraints and L2 norm regularization to prevent the instability of policy performance caused by too large updates. Finally, through the experimental verification of Decentralized Collective Assault (DCA), our algorithm has achieved significant improvements in various performance indicators, indicating that it has better effect and robustness in dealing with complex tasks.

源语言英语
主期刊名Proceedings of the 43rd Chinese Control Conference, CCC 2024
编辑Jing Na, Jian Sun
出版商IEEE Computer Society
6109-6114
页数6
ISBN(电子版)9789887581581
DOI
出版状态已出版 - 2024
活动43rd Chinese Control Conference, CCC 2024 - Kunming, 中国
期限: 28 7月 202431 7月 2024

出版系列

姓名Chinese Control Conference, CCC
ISSN(印刷版)1934-1768
ISSN(电子版)2161-2927

会议

会议43rd Chinese Control Conference, CCC 2024
国家/地区中国
Kunming
时期28/07/2431/07/24

指纹

探究 'Multi-time scale hierarchical trust domain leads to the improvement of MAPPO algorithm' 的科研主题。它们共同构成独一无二的指纹。

引用此