TY - GEN
T1 - A Graph-Based Reinforcement Learning Method for Flexible Job Shop Scheduling with Sequence Flexibility
AU - Li, Guohao
AU - Yuan, Erdong
AU - Wang, Liejun
AU - Song, Shiji
AU - Zhang, Yuli
AU - Zhong, Xiuxian
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2027.
PY - 2027
Y1 - 2027
N2 - The flexible job shop scheduling problem with sequence flexibility (FJSPSF) has a wide range of applications in various industries, including printing and semiconductor manufacturing. To effectively address the FJSPSF, we introduce a novel end-to-end deep reinforcement learning (DRL) framework. It is the first DRL method designed specifically to handle the challenges of sequence flexibility in this domain. Firstly, to capture the complex topological structure, which is given in the form of a directed acyclic graph (DAG), we design a DAG-based heterogeneous graph neural network (DAHGNN) to represent the shop state and obtain a high-quality state embedding. Based on the heterogeneous graph representation of scheduling states, the Proximal Policy Optimization (PPO) algorithm is adopted to learn an effective policy. Then, we propose a new Markov decision process (MDP) model for this problem, in which the state features are better designed and the action space is specially constructed for the characteristics of the FJSPSF. Finally, we conduct extensive experiments on two public datasets (DAFJS and YFJS), and the results demonstrate that the policy model outperforms all the priority dispatching rules (PDRs). Furthermore, it achieves solution qualities close to state-of-the-art metaheuristic and exact methods, while offering significantly faster solving speeds.
AB - The flexible job shop scheduling problem with sequence flexibility (FJSPSF) has a wide range of applications in various industries, including printing and semiconductor manufacturing. To effectively address the FJSPSF, we introduce a novel end-to-end deep reinforcement learning (DRL) framework. It is the first DRL method designed specifically to handle the challenges of sequence flexibility in this domain. Firstly, to capture the complex topological structure, which is given in the form of a directed acyclic graph (DAG), we design a DAG-based heterogeneous graph neural network (DAHGNN) to represent the shop state and obtain a high-quality state embedding. Based on the heterogeneous graph representation of scheduling states, the Proximal Policy Optimization (PPO) algorithm is adopted to learn an effective policy. Then, we propose a new Markov decision process (MDP) model for this problem, in which the state features are better designed and the action space is specially constructed for the characteristics of the FJSPSF. Finally, we conduct extensive experiments on two public datasets (DAFJS and YFJS), and the results demonstrate that the policy model outperforms all the priority dispatching rules (PDRs). Furthermore, it achieves solution qualities close to state-of-the-art metaheuristic and exact methods, while offering significantly faster solving speeds.
KW - Deep reinforcement learning
KW - Flexible job shop scheduling
KW - Graph neural network
KW - Sequence flexibility
UR - https://www.scopus.com/pages/publications/105046162491
U2 - 10.1007/978-981-92-3438-7_22
DO - 10.1007/978-981-92-3438-7_22
M3 - Conference contribution
AN - SCOPUS:105046162491
SN - 9789819234370
T3 - Lecture Notes in Computer Science
SP - 257
EP - 268
BT - Advanced Intelligent Computing Technology and Applications - 22nd International Conference on Intelligent Computing, ICIC 2026, Proceedings
A2 - Huang, De-Shuang
A2 - Zhang, Qinhu
A2 - Pan, Yijie
A2 - Zhang, Chuanlei
A2 - Chen, Wei
A2 - Li, Bo
A2 - Bao, Wenzheng
A2 - Premaratne, Prashan
PB - Springer Science and Business Media Deutschland GmbH
T2 - 22nd International Conference on Intelligent Computing, ICIC 2026
Y2 - 22 July 2026 through 26 July 2026
ER -