Abstract
This article presents a learning-based framework to approximate parameteric mixed-integer optimal control problem (MIOCP) solutions via mixed-integer differentiable predictive control (MI-DPC). To handle continuous and discrete controls, outer convexification and relaxed reformulation are applied. We further design the hierarchical neural network architecture and introduce the discrete postprocessing layer to represent them, guaranteeing the differentiability of discrete variables. However, relying solely on differentiable programming for solving MIOCP may easily lead to local optima. To mitigate it, we propose a mixed training architecture that combines group relative policy optimization with MI-DPC for policy enhancement. This approach leverages tailored policy exploration and group-relative comparison mechanisms to reinforce advantageous actions. Finally, the proposed method is applied to the numerical case and the energy management strategy of hybrid electric vehicles. Simulation and hardware-in-the-loop experiments demonstrate the real-time capability and effectiveness of the proposed framework. Compared with existing methods and mature solvers, the proposed approach improves policy optimality by 2%–4% and maintains submillisecond inference, providing an alternative for embedded system deployment.
| Original language | English |
|---|---|
| Journal | IEEE/ASME Transactions on Mechatronics |
| DOIs | |
| Publication status | Accepted/In press - 2026 |
Keywords
- Differentiable optimization
- hybrid electric vehicle (HEV)
- mixed-integer model predictive control (MIMPC)
- real-time solution
Fingerprint
Dive into the research topics of 'GRPO Enhanced Mixed-Integer Differentiable Predictive Control: A Case of Energy Management of HEV'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver