TY - JOUR
T1 - Dynamic Clustering Algorithm for UAV Ad Hoc Networks Based on Multi-Agent Reinforcement Learning
AU - Wu, Nan
AU - Wei, Shiyu
AU - Zhang, Tingting
AU - Fang, Jiarui
N1 - Publisher Copyright:
© (2026), (Chinese Institute of Electronics). All rights reserved.
PY - 2026
Y1 - 2026
N2 - With the development of swarm intelligence collaboration and network communication technologies, UAV ad hoc networks (UANETs) have demonstrated increasingly prominent advantages in mobility, flexibility, and envi⁃ ronmental adaptability, and have been widely applied in both military and civilian scenarios. However, due to the large net⁃ work scale, limited node energy, and highly dynamic topology of UANETs, traditional clustering algorithms often suffer from network instability and high energy consumption. To address these challenges, this paper proposes a dynamic cluster⁃ ing algorithm for UANETs based on multi-agent reinforcement learning (MARL). First, a clustering strategy based on ener⁃ gy and mobility constraints is designed. By jointly considering four factors, namely residual energy, node degree, inter-node distance, and link maintenance time, a dynamic weight calculation model tailored to the characteristics of UAV networks is constructed. During cluster formation, an evaluation mechanism for the clustering states of neighboring nodes is introduced. Each node performs cluster head election according to its own weight and the clustering states of its neighbors, and follows the cluster head based on link weights, thereby effectively reducing the number of isolated nodes. Furthermore, a distributed topology optimization mechanism based on independent Q-learning (IQL) is developed, where each node acts as an indepen⁃ dent agent and interacts with the dynamic network environment to evaluate the reward of following different cluster heads. When the link quality to the original cluster head degrades, each node autonomously optimizes its cluster affiliation decision according to the accumulated reward, thus enhancing the robustness of the cluster topology and achieving balanced energy consumption optimization across the network. Experimental results demonstrate that, compared with existing baseline algo⁃ rithms, the proposed algorithm achieves superior performance in terms of clustering efficiency, node energy consumption, and cluster stability, providing effective technical support for the reliable operation of UANETs in highly dynamic environ⁃ ments.
AB - With the development of swarm intelligence collaboration and network communication technologies, UAV ad hoc networks (UANETs) have demonstrated increasingly prominent advantages in mobility, flexibility, and envi⁃ ronmental adaptability, and have been widely applied in both military and civilian scenarios. However, due to the large net⁃ work scale, limited node energy, and highly dynamic topology of UANETs, traditional clustering algorithms often suffer from network instability and high energy consumption. To address these challenges, this paper proposes a dynamic cluster⁃ ing algorithm for UANETs based on multi-agent reinforcement learning (MARL). First, a clustering strategy based on ener⁃ gy and mobility constraints is designed. By jointly considering four factors, namely residual energy, node degree, inter-node distance, and link maintenance time, a dynamic weight calculation model tailored to the characteristics of UAV networks is constructed. During cluster formation, an evaluation mechanism for the clustering states of neighboring nodes is introduced. Each node performs cluster head election according to its own weight and the clustering states of its neighbors, and follows the cluster head based on link weights, thereby effectively reducing the number of isolated nodes. Furthermore, a distributed topology optimization mechanism based on independent Q-learning (IQL) is developed, where each node acts as an indepen⁃ dent agent and interacts with the dynamic network environment to evaluate the reward of following different cluster heads. When the link quality to the original cluster head degrades, each node autonomously optimizes its cluster affiliation decision according to the accumulated reward, thus enhancing the robustness of the cluster topology and achieving balanced energy consumption optimization across the network. Experimental results demonstrate that, compared with existing baseline algo⁃ rithms, the proposed algorithm achieves superior performance in terms of clustering efficiency, node energy consumption, and cluster stability, providing effective technical support for the reliable operation of UANETs in highly dynamic environ⁃ ments.
KW - cluster-head election
KW - dynamic clustering
KW - multi-agent reinforcement learning
KW - to⁃ pology optimization
KW - UAV ad hoc networks
KW - 动态分簇
KW - 多智能体强化学习
KW - 拓扑优化
KW - 无人机自组网
KW - 簇首选举
UR - https://www.scopus.com/pages/publications/105044214096
U2 - 10.12263/DZXB.20251147
DO - 10.12263/DZXB.20251147
M3 - Article
AN - SCOPUS:105044214096
SN - 0372-2112
VL - 54
SP - 1048
EP - 1061
JO - Tien Tzu Hsueh Pao/Acta Electronica Sinica
JF - Tien Tzu Hsueh Pao/Acta Electronica Sinica
IS - 3
ER -