Skip to main navigation Skip to search Skip to main content

Disturbance-Rejection Reinforcement Learning for Control of Unknown Nonlinear Systems

  • Beijing Institute of Technology
  • Zhejiang University of Technology
  • Zhongyuan University of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

How to realize the ultimate boundedness during training on unknown nonlinear systems under continuous stochastic disturbances is a key challenge in reinforcement learning (RL). To this end, we propose a Lyapunov-based data-driven RL framework for disturbance rejection control and provides a step toward reliable disturbance-rejection learning. First, a stabilizable reference model is constructed via input-convex neural networks (ICNNs) to generate stabilizable, disturbance-aware trajectories from data that reduce sensitivity to nominal model mismatch. Second, we establish a data-driven ultimate-boundedness theorem. It certifies closed-loop stability directly from sampled transitions without requiring prior model knowledge. Building on these foundations, we propose an off-policy algorithm that alternates between reference-model learning and policy optimization. It enforces Lyapunov decrease in expectation and jointly improves performance and robustness to achieve ultimate boundedness. Experiments on a classic control task and MuJoCo benchmarks demonstrate that the proposed method reduces the final-20-episode average cost by up to 75% and achieves average episode lengths over the final 20 episodes up to 3.2 times longer than those of representative Lyapunov-based RL methods.

Original languageEnglish
JournalIEEE Transactions on Automation Science and Engineering
DOIs
Publication statusAccepted/In press - 2026
Externally publishedYes

Keywords

  • Lyapunov stability
  • Reinforcement learning
  • robustness
  • ultimate boundedness

Fingerprint

Dive into the research topics of 'Disturbance-Rejection Reinforcement Learning for Control of Unknown Nonlinear Systems'. Together they form a unique fingerprint.

Cite this