Abstract
—Synergistically optimizing cooperative trajectory and task offloading in UAV-assisted Mobile Edge Computing (MEC) is a critical challenge for 6G-enabled ubiquitous computing. However, existing multi-agent reinforcement learning (MARL) frameworks suffer from severe gradient scale mismatch in hybrid action spaces, multiplier oscillations under non-convex fairness constraints, and strategic non-stationarity in dynamic environments. This paper proposes MAPA-MARL, a novel meta-policy framework designed to structurally resolve these bottlenecks. Specifically, we formulate the joint optimization as a Constrained Markov Decision Process (CMDP) and introduce a PID-Lagrangian mechanism to enforce Jain’s Fairness Index, utilizing Exponential Moving Average (EMA) filtered derivative damping to safely suppress the instability typical of traditional dual ascent. To handle the hybrid decision-making space, we develop a Dual-Head Actor architecture integrated with a Variance-Bounded Zero-temperature Gumbel-Rao (ZGR) estimator. This provides a low-variance empirical gradient proxy, effectively mitigating the gradient variance explosion and preventing the meta-initialization collapse induced by conventional discrete relaxations. Furthermore, a Sparsity-Induced Inverse-Distance Soft-Mask is proposed; rather than merely reducing computational complexity, it acts as a topological regularizer to block polluted gradients from distant agents, thereby stabilizing local coordination. To anchor convergence against peer-evolution without inter-UAV parameter sharing, we integrate an RNN-based implicit opponent modeling module augmented with Hessian-Vector Product (HVP) second-order curvature compensation. Extensive simulations demonstrate that MAPA-MARL achieves superior Pareto efficiency with a minimal 18.5 MB memory footprint and millisecond-level onboard inference latency, confirming its robust viability for resource-constrained edge deployment.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Vehicular Technology |
| DOIs | |
| Publication status | Accepted/In press - 2026 |
Keywords
- Meta-Learning
- Mobile Edge Computing (MEC)
- Multi-Agent Reinforcement Learning
- UAV
Fingerprint
Dive into the research topics of 'Meta-Policy-Based Multi-Agent Reinforcement Learning for Dynamic UAV Networks: Resolving Hybrid Action Spaces and Non-Stationarity'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver