Skip to main navigation Skip to search Skip to main content

Dual-Time-Scale Framework for Joint Optimization of Service Caching and UAV Trajectory Based on Self-Attention Deep Reinforcement Learning

  • Xuewei Zhang
  • , Jiantao Li
  • , Yuan Ren*
  • , Fan Jiang
  • , Junxuan Wang
  • , Jie Zeng
  • *Corresponding author for this work
  • Xi'an Institute of Posts and Telecommunications
  • Beijing Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Nowadays, unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) has been recognized as a promising technique for flexibly handling computation tasks in 5G advanced and 6G networks. This article investigates the joint optimization of service caching and computation offloading within a dual time-scale framework. We maximize the caching utility and minimize the task processing delay by jointly optimizing service caching policies, UAV flight trajectory, and computation offloading decisions. Specifically, for the long-term problem, we use the latent Dirichlet allocation (LDA) model to predict user preferences and propose a Lagrangian dual decomposition-based algorithm. For the short-term problem, a self-attention-based multiagent proximal policy optimization (MAPPO) algorithm is designed. Under the centralized training with decentralized execution (CTDE) framework, this algorithm integrates a multihead self-attention mechanism with curriculum learning. Each UAV is regarded as an agent, and a self-attention encoder (SAE) is integrated at the front-end of each actor network. This enables the agent to dynamically capture the relative importance between itself and all users, and context-aware features are extracted to make more intelligent and trajectory designs. Through extensive simulation experiments, the long-term algorithm yields the performance improvements of 76.5% in cache hit rate and 66% in caching utility, as compared to the second-best baseline. In dynamic scenarios, the short-term algorithm achieves a 16% reduction in total processing latency with respect to the proximal policy optimization (PPO) policy.

Original languageEnglish
Pages (from-to)34629-34645
Number of pages17
JournalIEEE Internet of Things Journal
Volume13
Issue number15
DOIs
Publication statusPublished - 1 Aug 2026
Externally publishedYes

Keywords

  • Mobile edge computing (MEC)
  • reinforcement learning
  • self-attention mechanism
  • unmanned aerial vehicle (UAV)

Fingerprint

Dive into the research topics of 'Dual-Time-Scale Framework for Joint Optimization of Service Caching and UAV Trajectory Based on Self-Attention Deep Reinforcement Learning'. Together they form a unique fingerprint.

Cite this