TY - JOUR
T1 - Dual-Time-Scale Framework for Joint Optimization of Service Caching and UAV Trajectory Based on Self-Attention Deep Reinforcement Learning
AU - Zhang, Xuewei
AU - Li, Jiantao
AU - Ren, Yuan
AU - Jiang, Fan
AU - Wang, Junxuan
AU - Zeng, Jie
N1 - Publisher Copyright:
© 2014 IEEE.
PY - 2026/8/1
Y1 - 2026/8/1
N2 - Nowadays, unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) has been recognized as a promising technique for flexibly handling computation tasks in 5G advanced and 6G networks. This article investigates the joint optimization of service caching and computation offloading within a dual time-scale framework. We maximize the caching utility and minimize the task processing delay by jointly optimizing service caching policies, UAV flight trajectory, and computation offloading decisions. Specifically, for the long-term problem, we use the latent Dirichlet allocation (LDA) model to predict user preferences and propose a Lagrangian dual decomposition-based algorithm. For the short-term problem, a self-attention-based multiagent proximal policy optimization (MAPPO) algorithm is designed. Under the centralized training with decentralized execution (CTDE) framework, this algorithm integrates a multihead self-attention mechanism with curriculum learning. Each UAV is regarded as an agent, and a self-attention encoder (SAE) is integrated at the front-end of each actor network. This enables the agent to dynamically capture the relative importance between itself and all users, and context-aware features are extracted to make more intelligent and trajectory designs. Through extensive simulation experiments, the long-term algorithm yields the performance improvements of 76.5% in cache hit rate and 66% in caching utility, as compared to the second-best baseline. In dynamic scenarios, the short-term algorithm achieves a 16% reduction in total processing latency with respect to the proximal policy optimization (PPO) policy.
AB - Nowadays, unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) has been recognized as a promising technique for flexibly handling computation tasks in 5G advanced and 6G networks. This article investigates the joint optimization of service caching and computation offloading within a dual time-scale framework. We maximize the caching utility and minimize the task processing delay by jointly optimizing service caching policies, UAV flight trajectory, and computation offloading decisions. Specifically, for the long-term problem, we use the latent Dirichlet allocation (LDA) model to predict user preferences and propose a Lagrangian dual decomposition-based algorithm. For the short-term problem, a self-attention-based multiagent proximal policy optimization (MAPPO) algorithm is designed. Under the centralized training with decentralized execution (CTDE) framework, this algorithm integrates a multihead self-attention mechanism with curriculum learning. Each UAV is regarded as an agent, and a self-attention encoder (SAE) is integrated at the front-end of each actor network. This enables the agent to dynamically capture the relative importance between itself and all users, and context-aware features are extracted to make more intelligent and trajectory designs. Through extensive simulation experiments, the long-term algorithm yields the performance improvements of 76.5% in cache hit rate and 66% in caching utility, as compared to the second-best baseline. In dynamic scenarios, the short-term algorithm achieves a 16% reduction in total processing latency with respect to the proximal policy optimization (PPO) policy.
KW - Mobile edge computing (MEC)
KW - reinforcement learning
KW - self-attention mechanism
KW - unmanned aerial vehicle (UAV)
UR - https://www.scopus.com/pages/publications/105040216441
U2 - 10.1109/JIOT.2026.3696305
DO - 10.1109/JIOT.2026.3696305
M3 - Article
AN - SCOPUS:105040216441
SN - 2327-4662
VL - 13
SP - 34629
EP - 34645
JO - IEEE Internet of Things Journal
JF - IEEE Internet of Things Journal
IS - 15
ER -