Dynamic flexible job-shop scheduling by multi-agent reinforcement learning with reward-shaping

Lixiang Zhang; Yan Yan; Chen Yang; Yaoguang Hu

doi:10.1016/j.aei.2024.102872

Dynamic flexible job-shop scheduling by multi-agent reinforcement learning with reward-shaping

Lixiang Zhang, Yan Yan, Chen Yang, Yaoguang Hu^*

^*Corresponding author for this work

Research output: Contribution to journal › Article › peer-review

5 Citations (Scopus)

Abstract

Achieving mass personalization presents significant challenges in performance and adaptability when solving dynamic flexible job-shop scheduling problems (DFJSP). Previous studies have struggled to achieve high performance in variable contexts. To tackle this challenge, this paper introduces a novel scheduling strategy founded on heterogeneous multi-agent reinforcement learning. This strategy facilitates centralized optimization and decentralized decision-making through collaboration among job and machine agents while employing historical experiences to support data-driven learning. The DFJSP with transportation time is initially formulated as heterogeneous multi-agent partial observation Markov Decision Processes. This formulation outlines the interactions between decision-making agents and the environment, incorporating a reward-shaping mechanism aimed at organizing job and machine agents to minimize the weighted tardiness of dynamic jobs. Then, we develop a dueling double deep Q-network algorithm incorporating the reward-shaping mechanism to ascertain the optimal strategies for machine allocation and job sequencing in DFJSP. This approach addresses the sparse reward issue and accelerates the learning process. Finally, the efficiency of the proposed method is verified and validated through numerical experiments, which demonstrate its superiority in reducing the weighted tardiness of dynamic jobs when compared to state-of-the-art baselines. The proposed method exhibits remarkable adaptability in encountering new scenarios, underscoring the benefits of adopting a heterogeneous multi-agent reinforcement learning-based scheduling approach in navigating dynamic and flexible challenges.

Original language	English
Article number	102872
Journal	Advanced Engineering Informatics
Volume	62
DOIs	https://doi.org/10.1016/j.aei.2024.102872
Publication status	Published - Oct 2024

Keywords

Deep reinforcement learning
Dynamic flexible job-shop scheduling
Multi-agent system
Reward-shaping

Access to Document

10.1016/j.aei.2024.102872

Cite this

Zhang, L., Yan, Y., Yang, C., & Hu, Y. (2024). Dynamic flexible job-shop scheduling by multi-agent reinforcement learning with reward-shaping. Advanced Engineering Informatics, 62, Article 102872. https://doi.org/10.1016/j.aei.2024.102872

@article{cf16c573233d48fea29ad219431ef5cd,

title = "Dynamic flexible job-shop scheduling by multi-agent reinforcement learning with reward-shaping",

abstract = "Achieving mass personalization presents significant challenges in performance and adaptability when solving dynamic flexible job-shop scheduling problems (DFJSP). Previous studies have struggled to achieve high performance in variable contexts. To tackle this challenge, this paper introduces a novel scheduling strategy founded on heterogeneous multi-agent reinforcement learning. This strategy facilitates centralized optimization and decentralized decision-making through collaboration among job and machine agents while employing historical experiences to support data-driven learning. The DFJSP with transportation time is initially formulated as heterogeneous multi-agent partial observation Markov Decision Processes. This formulation outlines the interactions between decision-making agents and the environment, incorporating a reward-shaping mechanism aimed at organizing job and machine agents to minimize the weighted tardiness of dynamic jobs. Then, we develop a dueling double deep Q-network algorithm incorporating the reward-shaping mechanism to ascertain the optimal strategies for machine allocation and job sequencing in DFJSP. This approach addresses the sparse reward issue and accelerates the learning process. Finally, the efficiency of the proposed method is verified and validated through numerical experiments, which demonstrate its superiority in reducing the weighted tardiness of dynamic jobs when compared to state-of-the-art baselines. The proposed method exhibits remarkable adaptability in encountering new scenarios, underscoring the benefits of adopting a heterogeneous multi-agent reinforcement learning-based scheduling approach in navigating dynamic and flexible challenges.",

keywords = "Deep reinforcement learning, Dynamic flexible job-shop scheduling, Multi-agent system, Reward-shaping",

author = "Lixiang Zhang and Yan Yan and Chen Yang and Yaoguang Hu",

note = "Publisher Copyright: {\textcopyright} 2024 Elsevier Ltd",

year = "2024",

month = oct,

doi = "10.1016/j.aei.2024.102872",

language = "English",

volume = "62",

journal = "Advanced Engineering Informatics",

issn = "1474-0346",

publisher = "Elsevier Ltd.",

}

TY - JOUR

T1 - Dynamic flexible job-shop scheduling by multi-agent reinforcement learning with reward-shaping

AU - Zhang, Lixiang

AU - Yan, Yan

AU - Yang, Chen

AU - Hu, Yaoguang

PY - 2024/10

Y1 - 2024/10

N2 - Achieving mass personalization presents significant challenges in performance and adaptability when solving dynamic flexible job-shop scheduling problems (DFJSP). Previous studies have struggled to achieve high performance in variable contexts. To tackle this challenge, this paper introduces a novel scheduling strategy founded on heterogeneous multi-agent reinforcement learning. This strategy facilitates centralized optimization and decentralized decision-making through collaboration among job and machine agents while employing historical experiences to support data-driven learning. The DFJSP with transportation time is initially formulated as heterogeneous multi-agent partial observation Markov Decision Processes. This formulation outlines the interactions between decision-making agents and the environment, incorporating a reward-shaping mechanism aimed at organizing job and machine agents to minimize the weighted tardiness of dynamic jobs. Then, we develop a dueling double deep Q-network algorithm incorporating the reward-shaping mechanism to ascertain the optimal strategies for machine allocation and job sequencing in DFJSP. This approach addresses the sparse reward issue and accelerates the learning process. Finally, the efficiency of the proposed method is verified and validated through numerical experiments, which demonstrate its superiority in reducing the weighted tardiness of dynamic jobs when compared to state-of-the-art baselines. The proposed method exhibits remarkable adaptability in encountering new scenarios, underscoring the benefits of adopting a heterogeneous multi-agent reinforcement learning-based scheduling approach in navigating dynamic and flexible challenges.

AB - Achieving mass personalization presents significant challenges in performance and adaptability when solving dynamic flexible job-shop scheduling problems (DFJSP). Previous studies have struggled to achieve high performance in variable contexts. To tackle this challenge, this paper introduces a novel scheduling strategy founded on heterogeneous multi-agent reinforcement learning. This strategy facilitates centralized optimization and decentralized decision-making through collaboration among job and machine agents while employing historical experiences to support data-driven learning. The DFJSP with transportation time is initially formulated as heterogeneous multi-agent partial observation Markov Decision Processes. This formulation outlines the interactions between decision-making agents and the environment, incorporating a reward-shaping mechanism aimed at organizing job and machine agents to minimize the weighted tardiness of dynamic jobs. Then, we develop a dueling double deep Q-network algorithm incorporating the reward-shaping mechanism to ascertain the optimal strategies for machine allocation and job sequencing in DFJSP. This approach addresses the sparse reward issue and accelerates the learning process. Finally, the efficiency of the proposed method is verified and validated through numerical experiments, which demonstrate its superiority in reducing the weighted tardiness of dynamic jobs when compared to state-of-the-art baselines. The proposed method exhibits remarkable adaptability in encountering new scenarios, underscoring the benefits of adopting a heterogeneous multi-agent reinforcement learning-based scheduling approach in navigating dynamic and flexible challenges.

KW - Deep reinforcement learning

KW - Dynamic flexible job-shop scheduling

KW - Multi-agent system

KW - Reward-shaping

UR - http://www.scopus.com/inward/record.url?scp=85206481679&partnerID=8YFLogxK

U2 - 10.1016/j.aei.2024.102872

DO - 10.1016/j.aei.2024.102872

M3 - Article

AN - SCOPUS:85206481679

SN - 1474-0346

VL - 62

JO - Advanced Engineering Informatics

JF - Advanced Engineering Informatics

M1 - 102872

ER -

Dynamic flexible job-shop scheduling by multi-agent reinforcement learning with reward-shaping

Abstract

Keywords

Access to Document

Other files and links

Fingerprint

Cite this