Abstract
How to realize the ultimate boundedness during training on unknown nonlinear systems under continuous stochastic disturbances is a key challenge in reinforcement learning (RL). To this end, we propose a Lyapunov-based data-driven RL framework for disturbance rejection control and provides a step toward reliable disturbance-rejection learning. First, a stabilizable reference model is constructed via input-convex neural networks (ICNNs) to generate stabilizable, disturbance-aware trajectories from data that reduce sensitivity to nominal model mismatch. Second, we establish a data-driven ultimate-boundedness theorem. It certifies closed-loop stability directly from sampled transitions without requiring prior model knowledge. Building on these foundations, we propose an off-policy algorithm that alternates between reference-model learning and policy optimization. It enforces Lyapunov decrease in expectation and jointly improves performance and robustness to achieve ultimate boundedness. Experiments on a classic control task and MuJoCo benchmarks demonstrate that the proposed method reduces the final-20-episode average cost by up to 75% and achieves average episode lengths over the final 20 episodes up to 3.2 times longer than those of representative Lyapunov-based RL methods.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Automation Science and Engineering |
| DOIs | |
| Publication status | Accepted/In press - 2026 |
| Externally published | Yes |
Keywords
- Lyapunov stability
- Reinforcement learning
- robustness
- ultimate boundedness
Fingerprint
Dive into the research topics of 'Disturbance-Rejection Reinforcement Learning for Control of Unknown Nonlinear Systems'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver