TY - JOUR
T1 - Adaptive Mode Switching in AoI-Aware Multi-UAV Hybrid MEC-DC Networks
T2 - A Multi-Agent Reinforcement Learning Approach
AU - Fu, Kang
AU - Zhao, Qingjie
N1 - Publisher Copyright:
© 2002-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Multi-UAV networks are promising for supporting time-sensitive IoT applications, yet most existing studies focus on a single service type and fail to address the coexistence of heterogeneous tasks with fundamentally different timeliness and resource characteristics. In hybrid MEC–DC systems, data collection (DC) and mobile edge computing (MEC) tasks exhibit distinct age of information (AoI) evolution rules and computation–communication couplings, which makes the AoI-aware joint optimization of trajectory planning, task scheduling, and service mode selection under energy, mobility, and communication constraints highly challenging. To tackle these challenges, we propose an adaptive mode-switching multi-agent reinforcement learning framework (AMS-MARL) based on heterogeneous-agent proximal policy optimization (HAPPO). Specifically, a randomized agent update order is employed to decompose the joint advantage into sequential individual advantages, enabling stable and decentralized learning. In addition, a rank-based adaptive reward shaping mechanism is designed to balance information freshness across heterogeneous sensor nodes (SNs) by adjusting reward weights based on AoI deviation from the global average. Extensive simulations under diverse spatial distributions, task ratios, and packet sizes show that AMS-MARL consistently outperforms state-of-the-art baselines in reducing AoI and exhibits strong robustness across varying system settings.
AB - Multi-UAV networks are promising for supporting time-sensitive IoT applications, yet most existing studies focus on a single service type and fail to address the coexistence of heterogeneous tasks with fundamentally different timeliness and resource characteristics. In hybrid MEC–DC systems, data collection (DC) and mobile edge computing (MEC) tasks exhibit distinct age of information (AoI) evolution rules and computation–communication couplings, which makes the AoI-aware joint optimization of trajectory planning, task scheduling, and service mode selection under energy, mobility, and communication constraints highly challenging. To tackle these challenges, we propose an adaptive mode-switching multi-agent reinforcement learning framework (AMS-MARL) based on heterogeneous-agent proximal policy optimization (HAPPO). Specifically, a randomized agent update order is employed to decompose the joint advantage into sequential individual advantages, enabling stable and decentralized learning. In addition, a rank-based adaptive reward shaping mechanism is designed to balance information freshness across heterogeneous sensor nodes (SNs) by adjusting reward weights based on AoI deviation from the global average. Extensive simulations under diverse spatial distributions, task ratios, and packet sizes show that AMS-MARL consistently outperforms state-of-the-art baselines in reducing AoI and exhibits strong robustness across varying system settings.
KW - Adaptive mode switching
KW - age of information (AoI)
KW - hybrid MEC–DC systems
KW - multi -agent reinforcement learning (MARL)
KW - multi -UAV networks
UR - https://www.scopus.com/pages/publications/105045553156
U2 - 10.1109/TMC.2026.3713132
DO - 10.1109/TMC.2026.3713132
M3 - Article
AN - SCOPUS:105045553156
SN - 1536-1233
JO - IEEE Transactions on Mobile Computing
JF - IEEE Transactions on Mobile Computing
ER -