Abstract
This paper proposes a data-efficient distributed adaptive dynamic programming (DADP) framework to address the challenge of time-varying formation control for unknown unmanned underwater vehicles (UUVs) under limited training data scenarios. First a distributed observer-based reference generator is designed, which decouples the exploration process of ADP for UUV from influencing neighboring UUVs’ learning outcomes. Second, to overcome data scarcity, we develop a dual-loop learning architecture: An inner-loop leverages least-squares (LS) iterations to find the LS solution from partial observations, reducing dependency on sufficient training data; An outer-loop uses policy iteration to optimize the control policy using Bellman equation. Theoretical analysis proves that the proposed DADP achieves convergence under limited training data. Comparing simulation results are provided to verify the advantages of the proposed approach over traditional ADP methods.
| Original language | English |
|---|---|
| Article number | 108820 |
| Journal | Journal of the Franklin Institute |
| Volume | 363 |
| Issue number | 12 |
| DOIs | |
| Publication status | Published - 1 Aug 2026 |
| Externally published | Yes |
Keywords
- Distributed adaptive dynamic programing
- Limited training data,
- Time-varying formation control
- Unmanned underwater vehicle
Fingerprint
Dive into the research topics of 'Data-efficient time-varying formation control of unknown unmanned underwater vehicles via distributed adaptive dynamic programming'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver