TY - JOUR
T1 - A Review of Reinforcement Learning for Multirotor UAVs from a Hierarchical Control Perspective
T2 - Biomimetic Architecture and Sim-to-Real
AU - Wei, Wei
AU - Zhao, Xubo
AU - Shu, Yongjie
AU - Meng, Qingkai
AU - Ding, Mingkai
AU - Wang, Yunyi
AU - Yan, Qingdong
N1 - Publisher Copyright:
© 2026 by the authors.
PY - 2026/6
Y1 - 2026/6
N2 - As unmanned aerial vehicle (UAV) systems evolve from automated execution toward autonomous decision-making, multirotor UAVs increasingly face complex dynamics, uncertain sensing conditions, and task-level autonomy demands. Reinforcement learning (RL) has emerged as a promising learning-based paradigm for addressing these challenges. Existing surveys on RL-based UAV control predominantly classify methods from an algorithmic or learning-paradigm perspective, while relatively little attention has been paid to the functional roles of RL policies within the control loop. This often leads to an unclear correspondence between algorithmic characteristics and the requirements of different control layers. To address this gap, this review proposes a biomimetic “spinal cord–cerebellum–cerebrum” framework, organizing existing RL studies into low-level dynamic stabilization, mid-level perception–action coordination, and high-level task planning and decision-making. The proposed hierarchy emphasizes the functional role and intervention depth of RL policies within the control architecture, further supporting a layer-wise analysis of sim-to-real challenges. This review aims to provide a structured understanding of the roles of reinforcement learning in hierarchical UAV control and to highlight future research directions toward robust real-world deployment.
AB - As unmanned aerial vehicle (UAV) systems evolve from automated execution toward autonomous decision-making, multirotor UAVs increasingly face complex dynamics, uncertain sensing conditions, and task-level autonomy demands. Reinforcement learning (RL) has emerged as a promising learning-based paradigm for addressing these challenges. Existing surveys on RL-based UAV control predominantly classify methods from an algorithmic or learning-paradigm perspective, while relatively little attention has been paid to the functional roles of RL policies within the control loop. This often leads to an unclear correspondence between algorithmic characteristics and the requirements of different control layers. To address this gap, this review proposes a biomimetic “spinal cord–cerebellum–cerebrum” framework, organizing existing RL studies into low-level dynamic stabilization, mid-level perception–action coordination, and high-level task planning and decision-making. The proposed hierarchy emphasizes the functional role and intervention depth of RL policies within the control architecture, further supporting a layer-wise analysis of sim-to-real challenges. This review aims to provide a structured understanding of the roles of reinforcement learning in hierarchical UAV control and to highlight future research directions toward robust real-world deployment.
KW - hierarchical control architecture
KW - multirotor UAVs
KW - reinforcement learning
KW - sim-to-real
UR - https://www.scopus.com/pages/publications/105042801897
U2 - 10.3390/drones10060448
DO - 10.3390/drones10060448
M3 - Review article
AN - SCOPUS:105042801897
SN - 2504-446X
VL - 10
JO - Drones
JF - Drones
IS - 6
M1 - 448
ER -