Skip to main navigation Skip to search Skip to main content

Al-Generated Expert Data Assisted Reinforcement Learning for Multi-agent Cooperative Tasks

  • Fenyu Guo
  • , Xinran Yi
  • , Yihan Ma
  • , Jiayi Hu
  • , Zhenjun Ming*
  • , Janet K. Allen
  • , Farrokh Mistree
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • University of Oklahoma

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

We can use multi-agent reinforcement learning algorithms such as Multi-Agent Deep Deterministic Policy Gradient (MADDPG) to effectively solve the organization problem of multi-agent systems. However, the training process of these algorithms is challenging due to sparse rewards, which means it typically takes a long time to finish the training. To address the challenge, we propose the Large Language Model (LLM) Assisted MADDPG (LLM-MADDPG) algorithm that utilizes LLM to guide the training of multi-agent systems. The core original contribution is two-fold. First is the dual-track mechanism, which combines the autonomous policy learning capability of the distributed Actor network with the cross-domain, high-level priori knowledge guidance provided by the Large Language Model (LLM). They both work together to improve training efficiency. The second is knowledge quality assurance through human-in-the-loop screening. The expert strategy knowledge generated by LLM is refined through manual screening and transformed into a code-based reward function to ensure logical rationality and accelerate the generation of correct guidance from LLM.

Original languageEnglish
Title of host publicationProceedings of 2025 9th Chinese Conference on Swarm Intelligence and Cooperative Control - Swarm Optimization Technologies
EditorsYongzhao Hua, Yishi Liu, Rui Yan
PublisherSpringer Science and Business Media Deutschland GmbH
Pages264-282
Number of pages19
ISBN (Print)9789819583287
DOIs
Publication statusPublished - 2026
Externally publishedYes
Event9th Chinese Conference on Swarm Intelligence and Cooperative Control, CCSICC 2025 - Shanghai, China
Duration: 31 Oct 20253 Nov 2025

Publication series

NameLecture Notes in Electrical Engineering
Volume1606 LNEE
ISSN (Print)1876-1100
ISSN (Electronic)1876-1119

Conference

Conference9th Chinese Conference on Swarm Intelligence and Cooperative Control, CCSICC 2025
Country/TerritoryChina
CityShanghai
Period31/10/253/11/25

Keywords

  • Hierarchical Reinforcement Learning
  • Large Language Model (LLM)
  • Multi-agent Reinforcement Learning

Fingerprint

Dive into the research topics of 'Al-Generated Expert Data Assisted Reinforcement Learning for Multi-agent Cooperative Tasks'. Together they form a unique fingerprint.

Cite this