TY - JOUR
T1 - Towards COLREGs-aware ship collision avoidance with multi-agent PPO-LSTM in maritime IoT
AU - Ding, Ying
AU - Meng, Weizhi
AU - He, Shaoming
AU - Li, Wenjuan
N1 - Publisher Copyright:
© 2026 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
PY - 2026/7
Y1 - 2026/7
N2 - Maritime Autonomous Surface Ships are expected to operate in a maritime IoT environment, where distributed sensing, V2V/AIS/VDES communication links, and electronic charts jointly support perception–decision–control loops for safe navigation in congested waters. A key challenge is to realise multi-ship collision avoidance that is consistent with the International Regulations for Preventing Collisions at Sea, while accounting for the limited manoeuvrability of large commercial vessels and the geometric constraints of ENC-derived chart-constrained narrow waterways. To address this problem, this work proposes a three-layer maritime IoT architecture in which each KVLCC2-class tanker is modelled as an IoT node, and ship states, TCPA/DCPA-based risk measures, and chart-derived environmental features are fused into a shared situational-awareness representation. On this basis, the task is formulated as a cooperative multi-agent partially observable Markov decision process, in which COLREGs encounter types, give-way/stand-on roles, and safety-domain constraints are embedded explicitly through the observation and reward design. A parameter-sharing recurrent multi-agent PPO–LSTM framework is then developed under the centralised-training-decentralised-execution paradigm, using a weakly centralised critic to handle partial observability and temporal coupling in dense multi-vessel interactions. The framework is evaluated in a unified simulation environment covering standard Imazu multi-vessel scenarios and an ENC-derived rasterised narrow-waterway case of Zhanjiang Bay, with comparisons against MA-PPO, MA-DDPG, and a classical VO baseline. Results show stronger convergence stability, higher mission success rates, larger closest-point-of-approach margins, and fewer COLREGs violations than the compared methods, while producing smooth and channel-conforming avoidance manoeuvres. Additional no-COLREG ablation and fixed-delay tests further clarify the roles of explicit rule-aware reward shaping and communication timeliness in cooperative multi-vessel collision avoidance.
AB - Maritime Autonomous Surface Ships are expected to operate in a maritime IoT environment, where distributed sensing, V2V/AIS/VDES communication links, and electronic charts jointly support perception–decision–control loops for safe navigation in congested waters. A key challenge is to realise multi-ship collision avoidance that is consistent with the International Regulations for Preventing Collisions at Sea, while accounting for the limited manoeuvrability of large commercial vessels and the geometric constraints of ENC-derived chart-constrained narrow waterways. To address this problem, this work proposes a three-layer maritime IoT architecture in which each KVLCC2-class tanker is modelled as an IoT node, and ship states, TCPA/DCPA-based risk measures, and chart-derived environmental features are fused into a shared situational-awareness representation. On this basis, the task is formulated as a cooperative multi-agent partially observable Markov decision process, in which COLREGs encounter types, give-way/stand-on roles, and safety-domain constraints are embedded explicitly through the observation and reward design. A parameter-sharing recurrent multi-agent PPO–LSTM framework is then developed under the centralised-training-decentralised-execution paradigm, using a weakly centralised critic to handle partial observability and temporal coupling in dense multi-vessel interactions. The framework is evaluated in a unified simulation environment covering standard Imazu multi-vessel scenarios and an ENC-derived rasterised narrow-waterway case of Zhanjiang Bay, with comparisons against MA-PPO, MA-DDPG, and a classical VO baseline. Results show stronger convergence stability, higher mission success rates, larger closest-point-of-approach margins, and fewer COLREGs violations than the compared methods, while producing smooth and channel-conforming avoidance manoeuvres. Additional no-COLREG ablation and fixed-delay tests further clarify the roles of explicit rule-aware reward shaping and communication timeliness in cooperative multi-vessel collision avoidance.
KW - Collision avoidance
KW - Digital twin
KW - Electronic Navigational Chart
KW - Maritime Internet of Things
KW - Multi-agent deep reinforcement learning
UR - https://www.scopus.com/pages/publications/105037662034
U2 - 10.1016/j.jnca.2026.104511
DO - 10.1016/j.jnca.2026.104511
M3 - Article
AN - SCOPUS:105037662034
SN - 1084-8045
VL - 251
JO - Journal of Network and Computer Applications
JF - Journal of Network and Computer Applications
M1 - 104511
ER -