Skip to main navigation Skip to search Skip to main content

From Semantic Decomposition to Human-Guided Control: Interpretable Multi-Agent Reinforcement Learning with Gated Expert

  • Beijing Institute of Technology
  • China Academy of Safety Science and Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Multi-Agent Reinforcement Learning (MARL) has shown remarkable performance in collaborative perception and decision-making tasks. However, its highly opaque decisionmaking architecture hinders human understanding, intervention, and correction of policy behavior, thereby limiting its application in safety-critical scenarios. This paper proposes an interpretable MARL framework based on Mixture-of-Experts (MoE) and human preference-guided gating, aiming to achieve interpretable decisions and controllable human-agent collaboration without compromising performance. Using a multi-agent target tracking task as the case study, the framework first pretrains a set of semantic experts and forms a frozen library of behavioral primitives. A learnable gating network is then introduced to dynamically fuse the experts' outputs through weighted aggregation, where the gating weights directly quantify the relative contribution of different semantic behaviors to the current decision, thereby providing intrinsic interpretability at the structural level. To better align the policy with human intent, we propose a human preference-guided gating adaptation mechanism that models human feedback as a form of structured supervision over the expert weight distribution, rather than direct modification of the reward function or policy parameters. This mechanism enables humans to adjust decision-making intent in a controllable manner without destabilizing the underlying behavioral primitives. Experimental results demonstrate that the proposed method significantly enhances training efficiency and policy stability in multi-agent target tracking tasks, achieving semantically consistent and intervenable decision-making behavior. This work offers a novel pathway for human-agent collaboration in interpretable reinforcement learning.

Original languageEnglish
Title of host publication38th Chinese Control and Decision Conference, CCDC 2026
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages7223-7228
Number of pages6
ISBN (Electronic)9798331550707
DOIs
Publication statusPublished - 2026
Externally publishedYes
Event38th Chinese Control and Decision Conference, CCDC 2026 - Nanjing, China
Duration: 15 May 202618 May 2026

Publication series

Name38th Chinese Control and Decision Conference, CCDC 2026

Conference

Conference38th Chinese Control and Decision Conference, CCDC 2026
Country/TerritoryChina
CityNanjing
Period15/05/2618/05/26

Keywords

  • Human-in-the-Loop Control
  • Interpretable Multi-Agent Reinforcement Learning
  • Mixture-of-Experts
  • Preference-Guided Policy Composition

Fingerprint

Dive into the research topics of 'From Semantic Decomposition to Human-Guided Control: Interpretable Multi-Agent Reinforcement Learning with Gated Expert'. Together they form a unique fingerprint.

Cite this