Skip to main navigation Skip to search Skip to main content

Data-efficient time-varying formation control of unknown unmanned underwater vehicles via distributed adaptive dynamic programming

  • Wanbing Zhao
  • , Ran Liu
  • , Jinliang Shao*
  • , Tieshan Li
  • , Yuhua Cheng
  • , Yuanqing Xia
  • *Corresponding author for this work
  • University of Electronic Science and Technology of China
  • Kunshan QTech Microelectronics Co.,Ltd.
  • Tianfu Jiangxi Laboratory
  • Zhongyuan University of Technology
  • Beijing Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

This paper proposes a data-efficient distributed adaptive dynamic programming (DADP) framework to address the challenge of time-varying formation control for unknown unmanned underwater vehicles (UUVs) under limited training data scenarios. First a distributed observer-based reference generator is designed, which decouples the exploration process of ADP for UUV from influencing neighboring UUVs’ learning outcomes. Second, to overcome data scarcity, we develop a dual-loop learning architecture: An inner-loop leverages least-squares (LS) iterations to find the LS solution from partial observations, reducing dependency on sufficient training data; An outer-loop uses policy iteration to optimize the control policy using Bellman equation. Theoretical analysis proves that the proposed DADP achieves convergence under limited training data. Comparing simulation results are provided to verify the advantages of the proposed approach over traditional ADP methods.

Original languageEnglish
Article number108820
JournalJournal of the Franklin Institute
Volume363
Issue number12
DOIs
Publication statusPublished - 1 Aug 2026
Externally publishedYes

Keywords

  • Distributed adaptive dynamic programing
  • Limited training data,
  • Time-varying formation control
  • Unmanned underwater vehicle

Fingerprint

Dive into the research topics of 'Data-efficient time-varying formation control of unknown unmanned underwater vehicles via distributed adaptive dynamic programming'. Together they form a unique fingerprint.

Cite this