Skip to main navigation Skip to search Skip to main content

Reinforcement learning-based trajectory planning for dual underwater manipulators

  • Zixuan Wang
  • , Shuyuan Pan
  • , Fangfei Cao*
  • *Corresponding author for this work
  • Beijing Institute of Technology

Research output: Contribution to journalConference articlepeer-review

Abstract

Dual underwater robotic manipulators are widely used in deep-sea exploration and subsea intervention tasks, where cooperative manipulation and environmental interaction impose significant challenges on trajectory planning. This paper investigates a reinforcement learning-based trajectory planning approach for dual underwater manipulators under kinematic constraints, collision avoidance requirements, and underwater disturbances. Proximal policy optimization (PPO) is employed to learn a continuous trajectory planning policy directly through interaction with the environment, benefiting from its stable training and robustness in high-dimensional control problems. To account for underwater operational characteristics, joint velocity is explicitly incorporated into the reward design to suppress aggressive motions, and mild flow disturbances are introduced during training to enhance robustness against environmental variability. Simulation results demonstrate that the proposed method can generate smooth, feasible, and coordinated trajectories for dual underwater manipulators under complex dynamics and underwater uncertainties.

Original languageEnglish
Pages (from-to)894-899
Number of pages6
JournalYouth Academic Annual Conference of Chinese Association of Automation, YAC
Issue number2026
DOIs
Publication statusPublished - 2026
Externally publishedYes
Event41st Youth Academic Annual Conference of Chinese Association of Automation, YAC 2026 - Changsha, China
Duration: 8 May 202610 May 2026

Keywords

  • dual underwater manipulators
  • hydrodynamic disturbance
  • reinforcement learning
  • trajectory planning

Fingerprint

Dive into the research topics of 'Reinforcement learning-based trajectory planning for dual underwater manipulators'. Together they form a unique fingerprint.

Cite this