TY - JOUR
T1 - GRPO Enhanced Mixed-Integer Differentiable Predictive Control
T2 - A Case of Energy Management of HEV
AU - Luo, Xi
AU - Dong, Shiying
AU - Hong, Jinlong
AU - Sun, Jian
AU - Yu, Haiyang
AU - Gao, Bingzhao
AU - Chen, Hong
N1 - Publisher Copyright:
© 1996-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - This article presents a learning-based framework to approximate parameteric mixed-integer optimal control problem (MIOCP) solutions via mixed-integer differentiable predictive control (MI-DPC). To handle continuous and discrete controls, outer convexification and relaxed reformulation are applied. We further design the hierarchical neural network architecture and introduce the discrete postprocessing layer to represent them, guaranteeing the differentiability of discrete variables. However, relying solely on differentiable programming for solving MIOCP may easily lead to local optima. To mitigate it, we propose a mixed training architecture that combines group relative policy optimization with MI-DPC for policy enhancement. This approach leverages tailored policy exploration and group-relative comparison mechanisms to reinforce advantageous actions. Finally, the proposed method is applied to the numerical case and the energy management strategy of hybrid electric vehicles. Simulation and hardware-in-the-loop experiments demonstrate the real-time capability and effectiveness of the proposed framework. Compared with existing methods and mature solvers, the proposed approach improves policy optimality by 2%–4% and maintains submillisecond inference, providing an alternative for embedded system deployment.
AB - This article presents a learning-based framework to approximate parameteric mixed-integer optimal control problem (MIOCP) solutions via mixed-integer differentiable predictive control (MI-DPC). To handle continuous and discrete controls, outer convexification and relaxed reformulation are applied. We further design the hierarchical neural network architecture and introduce the discrete postprocessing layer to represent them, guaranteeing the differentiability of discrete variables. However, relying solely on differentiable programming for solving MIOCP may easily lead to local optima. To mitigate it, we propose a mixed training architecture that combines group relative policy optimization with MI-DPC for policy enhancement. This approach leverages tailored policy exploration and group-relative comparison mechanisms to reinforce advantageous actions. Finally, the proposed method is applied to the numerical case and the energy management strategy of hybrid electric vehicles. Simulation and hardware-in-the-loop experiments demonstrate the real-time capability and effectiveness of the proposed framework. Compared with existing methods and mature solvers, the proposed approach improves policy optimality by 2%–4% and maintains submillisecond inference, providing an alternative for embedded system deployment.
KW - Differentiable optimization
KW - hybrid electric vehicle (HEV)
KW - mixed-integer model predictive control (MIMPC)
KW - real-time solution
UR - https://www.scopus.com/pages/publications/105044014165
U2 - 10.1109/TMECH.2026.3704664
DO - 10.1109/TMECH.2026.3704664
M3 - Article
AN - SCOPUS:105044014165
SN - 1083-4435
JO - IEEE/ASME Transactions on Mechatronics
JF - IEEE/ASME Transactions on Mechatronics
ER -