TY - JOUR
T1 - A Bi-level Deep Reinforcement Learning Framework for Task Allocation with Tight Constraints in Heterogeneous Multi-Agent Systems
AU - Hu, Jiatao
AU - Peng, Xiuhui
AU - Lv, Yuezu
AU - Zheng, Shaoqiu
N1 - Publisher Copyright:
© 1967-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - This paper addresses tightly constrained heterogeneous multi-agent task allocation (H-MATA) in physical-space missions, where agents must form role-complete coalitions and execute tasks under synchronization and execution-feasibility constraints. Existing optimization-based methods struggle with scalability under coupled coalition and execution constraints, whereas purely end-to-end MARL methods often suffer from inefficient exploration because feasible joint actions are sparse in the combinatorial action space. We propose the Deep Reinforcement Learning-based Heterogeneous Multi-agent Collaborative Mission Planning (DRL-HMAMP), a bi-level framework that decouples combinatorial coalition feasibility from continuous execution feedback while preserving closed-loop coupling between task selection and motion-level cost. The actor produces role-aware task intentions, the Dynamic Team Coordination System (DTCS) provides bounded coalition-feasibility guidance, and the Path Planner with Feedback (PPF) executes low-level motions while feeding path-cost information back into subsequent decisions. Evaluated over five independent random seeds, DRL-HMAMP achieves an average episode reward of 159.67, a completion rate of 0.95, and an average episode length of 63.43, yielding a 9.74% reward increase, a 4.40% completion-rate increase, and an 18.00% reduction in episode length compared with strong baselines.
AB - This paper addresses tightly constrained heterogeneous multi-agent task allocation (H-MATA) in physical-space missions, where agents must form role-complete coalitions and execute tasks under synchronization and execution-feasibility constraints. Existing optimization-based methods struggle with scalability under coupled coalition and execution constraints, whereas purely end-to-end MARL methods often suffer from inefficient exploration because feasible joint actions are sparse in the combinatorial action space. We propose the Deep Reinforcement Learning-based Heterogeneous Multi-agent Collaborative Mission Planning (DRL-HMAMP), a bi-level framework that decouples combinatorial coalition feasibility from continuous execution feedback while preserving closed-loop coupling between task selection and motion-level cost. The actor produces role-aware task intentions, the Dynamic Team Coordination System (DTCS) provides bounded coalition-feasibility guidance, and the Path Planner with Feedback (PPF) executes low-level motions while feeding path-cost information back into subsequent decisions. Evaluated over five independent random seeds, DRL-HMAMP achieves an average episode reward of 159.67, a completion rate of 0.95, and an average episode length of 63.43, yielding a 9.74% reward increase, a 4.40% completion-rate increase, and an 18.00% reduction in episode length compared with strong baselines.
KW - Coalition formation
KW - Multi-robot system
KW - Task allocation
KW - multi-agent reinforcement learning
UR - https://www.scopus.com/pages/publications/105041470972
U2 - 10.1109/TVT.2026.3701968
DO - 10.1109/TVT.2026.3701968
M3 - Article
AN - SCOPUS:105041470972
SN - 0018-9545
JO - IEEE Transactions on Vehicular Technology
JF - IEEE Transactions on Vehicular Technology
ER -