TY - GEN
T1 - Al-Generated Expert Data Assisted Reinforcement Learning for Multi-agent Cooperative Tasks
AU - Guo, Fenyu
AU - Yi, Xinran
AU - Ma, Yihan
AU - Hu, Jiayi
AU - Ming, Zhenjun
AU - Allen, Janet K.
AU - Mistree, Farrokh
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - We can use multi-agent reinforcement learning algorithms such as Multi-Agent Deep Deterministic Policy Gradient (MADDPG) to effectively solve the organization problem of multi-agent systems. However, the training process of these algorithms is challenging due to sparse rewards, which means it typically takes a long time to finish the training. To address the challenge, we propose the Large Language Model (LLM) Assisted MADDPG (LLM-MADDPG) algorithm that utilizes LLM to guide the training of multi-agent systems. The core original contribution is two-fold. First is the dual-track mechanism, which combines the autonomous policy learning capability of the distributed Actor network with the cross-domain, high-level priori knowledge guidance provided by the Large Language Model (LLM). They both work together to improve training efficiency. The second is knowledge quality assurance through human-in-the-loop screening. The expert strategy knowledge generated by LLM is refined through manual screening and transformed into a code-based reward function to ensure logical rationality and accelerate the generation of correct guidance from LLM.
AB - We can use multi-agent reinforcement learning algorithms such as Multi-Agent Deep Deterministic Policy Gradient (MADDPG) to effectively solve the organization problem of multi-agent systems. However, the training process of these algorithms is challenging due to sparse rewards, which means it typically takes a long time to finish the training. To address the challenge, we propose the Large Language Model (LLM) Assisted MADDPG (LLM-MADDPG) algorithm that utilizes LLM to guide the training of multi-agent systems. The core original contribution is two-fold. First is the dual-track mechanism, which combines the autonomous policy learning capability of the distributed Actor network with the cross-domain, high-level priori knowledge guidance provided by the Large Language Model (LLM). They both work together to improve training efficiency. The second is knowledge quality assurance through human-in-the-loop screening. The expert strategy knowledge generated by LLM is refined through manual screening and transformed into a code-based reward function to ensure logical rationality and accelerate the generation of correct guidance from LLM.
KW - Hierarchical Reinforcement Learning
KW - Large Language Model (LLM)
KW - Multi-agent Reinforcement Learning
UR - https://www.scopus.com/pages/publications/105040543366
U2 - 10.1007/978-981-95-8329-4_22
DO - 10.1007/978-981-95-8329-4_22
M3 - Conference contribution
AN - SCOPUS:105040543366
SN - 9789819583287
T3 - Lecture Notes in Electrical Engineering
SP - 264
EP - 282
BT - Proceedings of 2025 9th Chinese Conference on Swarm Intelligence and Cooperative Control - Swarm Optimization Technologies
A2 - Hua, Yongzhao
A2 - Liu, Yishi
A2 - Yan, Rui
PB - Springer Science and Business Media Deutschland GmbH
T2 - 9th Chinese Conference on Swarm Intelligence and Cooperative Control, CCSICC 2025
Y2 - 31 October 2025 through 3 November 2025
ER -