TY - GEN
T1 - Incremental Safe Reinforcement Learning for Flight Control with Control Barrier Function
AU - Guo, Yuze
AU - Liu, Junhui
AU - Wang, Jianan
AU - Shan, Jiayuan
AU - Zhang, Huan
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - In this paper, an incremental safe reinforcement learning algorithm, namely Incremental Safe Dual Heuristic Programming (ISDHP), is proposed for flight control systems with state constraints. Firstly, a recursive least squares (RLS) method is employed for online identification of incremental system dynamics to achieve model-free real-time adaptation without offline training. Secondly, the cost function is augmented with the control barrier function (CBF) to ensure that state constraints are satisfied. Thirdly, an actor-critic network structure with experience replay is designed to approximate the optimal control policy, where network weights are updated via gradient descent to minimize the temporal difference error. Finally, numerical simulations on a aircraft longitudinal dynamic model validate that the proposed ISDHP algorithm achieves effective tracking of the reference command while strictly confining the angle of attack within the safe range, demonstrating its superiority in both optimality and safety.
AB - In this paper, an incremental safe reinforcement learning algorithm, namely Incremental Safe Dual Heuristic Programming (ISDHP), is proposed for flight control systems with state constraints. Firstly, a recursive least squares (RLS) method is employed for online identification of incremental system dynamics to achieve model-free real-time adaptation without offline training. Secondly, the cost function is augmented with the control barrier function (CBF) to ensure that state constraints are satisfied. Thirdly, an actor-critic network structure with experience replay is designed to approximate the optimal control policy, where network weights are updated via gradient descent to minimize the temporal difference error. Finally, numerical simulations on a aircraft longitudinal dynamic model validate that the proposed ISDHP algorithm achieves effective tracking of the reference command while strictly confining the angle of attack within the safe range, demonstrating its superiority in both optimality and safety.
KW - control barrier function
KW - dual heuristic programming
KW - flight control
KW - online learning
KW - safe reinforcement learning
UR - https://www.scopus.com/pages/publications/105042893102
U2 - 10.1007/978-981-95-8435-2_47
DO - 10.1007/978-981-95-8435-2_47
M3 - Conference contribution
AN - SCOPUS:105042893102
SN - 9789819584345
T3 - Lecture Notes in Electrical Engineering
SP - 574
EP - 587
BT - Proceedings of 2025 9th Chinese Conference on Swarm Intelligence and Cooperative Control - Swarm Control Technologies
A2 - Wang, Qing
A2 - Dong, Xiwang
A2 - Song, Peng
PB - Springer Science and Business Media Deutschland GmbH
T2 - 9th Chinese Conference on Swarm Intelligence and Cooperative Control, CCSICC 2025
Y2 - 31 October 2025 through 3 November 2025
ER -