Skip to main navigation Skip to search Skip to main content

A Review of Reinforcement Learning for Multirotor UAVs from a Hierarchical Control Perspective: Biomimetic Architecture and Sim-to-Real

  • Wei Wei
  • , Xubo Zhao
  • , Yongjie Shu*
  • , Qingkai Meng
  • , Mingkai Ding
  • , Yunyi Wang
  • , Qingdong Yan
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • B&H Unmanned Intelligent System Research Institute

Research output: Contribution to journalReview articlepeer-review

Abstract

As unmanned aerial vehicle (UAV) systems evolve from automated execution toward autonomous decision-making, multirotor UAVs increasingly face complex dynamics, uncertain sensing conditions, and task-level autonomy demands. Reinforcement learning (RL) has emerged as a promising learning-based paradigm for addressing these challenges. Existing surveys on RL-based UAV control predominantly classify methods from an algorithmic or learning-paradigm perspective, while relatively little attention has been paid to the functional roles of RL policies within the control loop. This often leads to an unclear correspondence between algorithmic characteristics and the requirements of different control layers. To address this gap, this review proposes a biomimetic “spinal cord–cerebellum–cerebrum” framework, organizing existing RL studies into low-level dynamic stabilization, mid-level perception–action coordination, and high-level task planning and decision-making. The proposed hierarchy emphasizes the functional role and intervention depth of RL policies within the control architecture, further supporting a layer-wise analysis of sim-to-real challenges. This review aims to provide a structured understanding of the roles of reinforcement learning in hierarchical UAV control and to highlight future research directions toward robust real-world deployment.

Original languageEnglish
Article number448
JournalDrones
Volume10
Issue number6
DOIs
Publication statusPublished - Jun 2026

Keywords

  • hierarchical control architecture
  • multirotor UAVs
  • reinforcement learning
  • sim-to-real

Fingerprint

Dive into the research topics of 'A Review of Reinforcement Learning for Multirotor UAVs from a Hierarchical Control Perspective: Biomimetic Architecture and Sim-to-Real'. Together they form a unique fingerprint.

Cite this