跳到主要导航 跳到搜索 跳到主要内容

Meta-Policy-Based Multi-Agent Reinforcement Learning for Dynamic UAV Networks: Resolving Hybrid Action Spaces and Non-Stationarity

  • Zifeng Dai
  • , Hui Xie
  • , Shengjun Wei*
  • , Changzhen Hu
  • *此作品的通讯作者
  • Beijing Institute of Technology

科研成果: 期刊稿件文章同行评审

摘要

—Synergistically optimizing cooperative trajectory and task offloading in UAV-assisted Mobile Edge Computing (MEC) is a critical challenge for 6G-enabled ubiquitous computing. However, existing multi-agent reinforcement learning (MARL) frameworks suffer from severe gradient scale mismatch in hybrid action spaces, multiplier oscillations under non-convex fairness constraints, and strategic non-stationarity in dynamic environments. This paper proposes MAPA-MARL, a novel meta-policy framework designed to structurally resolve these bottlenecks. Specifically, we formulate the joint optimization as a Constrained Markov Decision Process (CMDP) and introduce a PID-Lagrangian mechanism to enforce Jain’s Fairness Index, utilizing Exponential Moving Average (EMA) filtered derivative damping to safely suppress the instability typical of traditional dual ascent. To handle the hybrid decision-making space, we develop a Dual-Head Actor architecture integrated with a Variance-Bounded Zero-temperature Gumbel-Rao (ZGR) estimator. This provides a low-variance empirical gradient proxy, effectively mitigating the gradient variance explosion and preventing the meta-initialization collapse induced by conventional discrete relaxations. Furthermore, a Sparsity-Induced Inverse-Distance Soft-Mask is proposed; rather than merely reducing computational complexity, it acts as a topological regularizer to block polluted gradients from distant agents, thereby stabilizing local coordination. To anchor convergence against peer-evolution without inter-UAV parameter sharing, we integrate an RNN-based implicit opponent modeling module augmented with Hessian-Vector Product (HVP) second-order curvature compensation. Extensive simulations demonstrate that MAPA-MARL achieves superior Pareto efficiency with a minimal 18.5 MB memory footprint and millisecond-level onboard inference latency, confirming its robust viability for resource-constrained edge deployment.

源语言英语
期刊IEEE Transactions on Vehicular Technology
DOI
出版状态已接受/待刊 - 2026

指纹

探究 'Meta-Policy-Based Multi-Agent Reinforcement Learning for Dynamic UAV Networks: Resolving Hybrid Action Spaces and Non-Stationarity' 的科研主题。它们共同构成独一无二的指纹。

引用此