TY - JOUR
T1 - Meta-Policy-Based Multi-Agent Reinforcement Learning for Dynamic UAV Networks
T2 - Resolving Hybrid Action Spaces and Non-Stationarity
AU - Dai, Zifeng
AU - Xie, Hui
AU - Wei, Shengjun
AU - Hu, Changzhen
N1 - Publisher Copyright:
© 2026 IEEE. All rights reserved,
PY - 2026
Y1 - 2026
N2 - —Synergistically optimizing cooperative trajectory and task offloading in UAV-assisted Mobile Edge Computing (MEC) is a critical challenge for 6G-enabled ubiquitous computing. However, existing multi-agent reinforcement learning (MARL) frameworks suffer from severe gradient scale mismatch in hybrid action spaces, multiplier oscillations under non-convex fairness constraints, and strategic non-stationarity in dynamic environments. This paper proposes MAPA-MARL, a novel meta-policy framework designed to structurally resolve these bottlenecks. Specifically, we formulate the joint optimization as a Constrained Markov Decision Process (CMDP) and introduce a PID-Lagrangian mechanism to enforce Jain’s Fairness Index, utilizing Exponential Moving Average (EMA) filtered derivative damping to safely suppress the instability typical of traditional dual ascent. To handle the hybrid decision-making space, we develop a Dual-Head Actor architecture integrated with a Variance-Bounded Zero-temperature Gumbel-Rao (ZGR) estimator. This provides a low-variance empirical gradient proxy, effectively mitigating the gradient variance explosion and preventing the meta-initialization collapse induced by conventional discrete relaxations. Furthermore, a Sparsity-Induced Inverse-Distance Soft-Mask is proposed; rather than merely reducing computational complexity, it acts as a topological regularizer to block polluted gradients from distant agents, thereby stabilizing local coordination. To anchor convergence against peer-evolution without inter-UAV parameter sharing, we integrate an RNN-based implicit opponent modeling module augmented with Hessian-Vector Product (HVP) second-order curvature compensation. Extensive simulations demonstrate that MAPA-MARL achieves superior Pareto efficiency with a minimal 18.5 MB memory footprint and millisecond-level onboard inference latency, confirming its robust viability for resource-constrained edge deployment.
AB - —Synergistically optimizing cooperative trajectory and task offloading in UAV-assisted Mobile Edge Computing (MEC) is a critical challenge for 6G-enabled ubiquitous computing. However, existing multi-agent reinforcement learning (MARL) frameworks suffer from severe gradient scale mismatch in hybrid action spaces, multiplier oscillations under non-convex fairness constraints, and strategic non-stationarity in dynamic environments. This paper proposes MAPA-MARL, a novel meta-policy framework designed to structurally resolve these bottlenecks. Specifically, we formulate the joint optimization as a Constrained Markov Decision Process (CMDP) and introduce a PID-Lagrangian mechanism to enforce Jain’s Fairness Index, utilizing Exponential Moving Average (EMA) filtered derivative damping to safely suppress the instability typical of traditional dual ascent. To handle the hybrid decision-making space, we develop a Dual-Head Actor architecture integrated with a Variance-Bounded Zero-temperature Gumbel-Rao (ZGR) estimator. This provides a low-variance empirical gradient proxy, effectively mitigating the gradient variance explosion and preventing the meta-initialization collapse induced by conventional discrete relaxations. Furthermore, a Sparsity-Induced Inverse-Distance Soft-Mask is proposed; rather than merely reducing computational complexity, it acts as a topological regularizer to block polluted gradients from distant agents, thereby stabilizing local coordination. To anchor convergence against peer-evolution without inter-UAV parameter sharing, we integrate an RNN-based implicit opponent modeling module augmented with Hessian-Vector Product (HVP) second-order curvature compensation. Extensive simulations demonstrate that MAPA-MARL achieves superior Pareto efficiency with a minimal 18.5 MB memory footprint and millisecond-level onboard inference latency, confirming its robust viability for resource-constrained edge deployment.
KW - Meta-Learning
KW - Mobile Edge Computing (MEC)
KW - Multi-Agent Reinforcement Learning
KW - UAV
UR - https://www.scopus.com/pages/publications/105041259280
U2 - 10.1109/TVT.2026.3700721
DO - 10.1109/TVT.2026.3700721
M3 - Article
AN - SCOPUS:105041259280
SN - 0018-9545
JO - IEEE Transactions on Vehicular Technology
JF - IEEE Transactions on Vehicular Technology
ER -