TY - JOUR
T1 - Disturbance-Rejection Reinforcement Learning for Control of Unknown Nonlinear Systems
AU - Gong, Hengheng
AU - Yan, Chenhang
AU - Yan, Tijin
AU - Zhan, Yufeng
AU - Xia, Yuanqing
N1 - Publisher Copyright:
© 2026 IEEE. All rights reserved.
PY - 2026
Y1 - 2026
N2 - How to realize the ultimate boundedness during training on unknown nonlinear systems under continuous stochastic disturbances is a key challenge in reinforcement learning (RL). To this end, we propose a Lyapunov-based data-driven RL framework for disturbance rejection control and provides a step toward reliable disturbance-rejection learning. First, a stabilizable reference model is constructed via input-convex neural networks (ICNNs) to generate stabilizable, disturbance-aware trajectories from data that reduce sensitivity to nominal model mismatch. Second, we establish a data-driven ultimate-boundedness theorem. It certifies closed-loop stability directly from sampled transitions without requiring prior model knowledge. Building on these foundations, we propose an off-policy algorithm that alternates between reference-model learning and policy optimization. It enforces Lyapunov decrease in expectation and jointly improves performance and robustness to achieve ultimate boundedness. Experiments on a classic control task and MuJoCo benchmarks demonstrate that the proposed method reduces the final-20-episode average cost by up to 75% and achieves average episode lengths over the final 20 episodes up to 3.2 times longer than those of representative Lyapunov-based RL methods. the framework guarantees bounded state trajectories, ensuring safety guarantees while maintaining performance under unknown conditions.
AB - How to realize the ultimate boundedness during training on unknown nonlinear systems under continuous stochastic disturbances is a key challenge in reinforcement learning (RL). To this end, we propose a Lyapunov-based data-driven RL framework for disturbance rejection control and provides a step toward reliable disturbance-rejection learning. First, a stabilizable reference model is constructed via input-convex neural networks (ICNNs) to generate stabilizable, disturbance-aware trajectories from data that reduce sensitivity to nominal model mismatch. Second, we establish a data-driven ultimate-boundedness theorem. It certifies closed-loop stability directly from sampled transitions without requiring prior model knowledge. Building on these foundations, we propose an off-policy algorithm that alternates between reference-model learning and policy optimization. It enforces Lyapunov decrease in expectation and jointly improves performance and robustness to achieve ultimate boundedness. Experiments on a classic control task and MuJoCo benchmarks demonstrate that the proposed method reduces the final-20-episode average cost by up to 75% and achieves average episode lengths over the final 20 episodes up to 3.2 times longer than those of representative Lyapunov-based RL methods. the framework guarantees bounded state trajectories, ensuring safety guarantees while maintaining performance under unknown conditions.
KW - Lyapunov stability
KW - Reinforcement learning
KW - robustness
KW - ultimate boundedness
UR - https://www.scopus.com/pages/publications/105046331145
U2 - 10.1109/TASE.2026.3718307
DO - 10.1109/TASE.2026.3718307
M3 - Article
AN - SCOPUS:105046331145
SN - 1545-5955
VL - 23
SP - 15238
EP - 15255
JO - IEEE Transactions on Automation Science and Engineering
JF - IEEE Transactions on Automation Science and Engineering
ER -