跳到主要导航 跳到搜索 跳到主要内容

Al-Generated Expert Data Assisted Reinforcement Learning for Multi-agent Cooperative Tasks

  • Fenyu Guo
  • , Xinran Yi
  • , Yihan Ma
  • , Jiayi Hu
  • , Zhenjun Ming*
  • , Janet K. Allen
  • , Farrokh Mistree
  • *此作品的通讯作者
  • Beijing Institute of Technology
  • University of Oklahoma

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

We can use multi-agent reinforcement learning algorithms such as Multi-Agent Deep Deterministic Policy Gradient (MADDPG) to effectively solve the organization problem of multi-agent systems. However, the training process of these algorithms is challenging due to sparse rewards, which means it typically takes a long time to finish the training. To address the challenge, we propose the Large Language Model (LLM) Assisted MADDPG (LLM-MADDPG) algorithm that utilizes LLM to guide the training of multi-agent systems. The core original contribution is two-fold. First is the dual-track mechanism, which combines the autonomous policy learning capability of the distributed Actor network with the cross-domain, high-level priori knowledge guidance provided by the Large Language Model (LLM). They both work together to improve training efficiency. The second is knowledge quality assurance through human-in-the-loop screening. The expert strategy knowledge generated by LLM is refined through manual screening and transformed into a code-based reward function to ensure logical rationality and accelerate the generation of correct guidance from LLM.

源语言英语
主期刊名Proceedings of 2025 9th Chinese Conference on Swarm Intelligence and Cooperative Control - Swarm Optimization Technologies
编辑Yongzhao Hua, Yishi Liu, Rui Yan
出版商Springer Science and Business Media Deutschland GmbH
264-282
页数19
ISBN(印刷版)9789819583287
DOI
出版状态已出版 - 2026
已对外发布
活动9th Chinese Conference on Swarm Intelligence and Cooperative Control, CCSICC 2025 - Shanghai, 中国
期限: 31 10月 20253 11月 2025

出版系列

姓名Lecture Notes in Electrical Engineering
1606 LNEE
ISSN(印刷版)1876-1100
ISSN(电子版)1876-1119

会议

会议9th Chinese Conference on Swarm Intelligence and Cooperative Control, CCSICC 2025
国家/地区中国
Shanghai
时期31/10/253/11/25

指纹

探究 'Al-Generated Expert Data Assisted Reinforcement Learning for Multi-agent Cooperative Tasks' 的科研主题。它们共同构成独一无二的指纹。

引用此