TY - JOUR
T1 - Data-efficient time-varying formation control of unknown unmanned underwater vehicles via distributed adaptive dynamic programming
AU - Zhao, Wanbing
AU - Liu, Ran
AU - Shao, Jinliang
AU - Li, Tieshan
AU - Cheng, Yuhua
AU - Xia, Yuanqing
N1 - Publisher Copyright:
© 2026 Published by Elsevier Inc. on behalf of The Franklin Institute.
PY - 2026/8/1
Y1 - 2026/8/1
N2 - This paper proposes a data-efficient distributed adaptive dynamic programming (DADP) framework to address the challenge of time-varying formation control for unknown unmanned underwater vehicles (UUVs) under limited training data scenarios. First a distributed observer-based reference generator is designed, which decouples the exploration process of ADP for UUV from influencing neighboring UUVs’ learning outcomes. Second, to overcome data scarcity, we develop a dual-loop learning architecture: An inner-loop leverages least-squares (LS) iterations to find the LS solution from partial observations, reducing dependency on sufficient training data; An outer-loop uses policy iteration to optimize the control policy using Bellman equation. Theoretical analysis proves that the proposed DADP achieves convergence under limited training data. Comparing simulation results are provided to verify the advantages of the proposed approach over traditional ADP methods.
AB - This paper proposes a data-efficient distributed adaptive dynamic programming (DADP) framework to address the challenge of time-varying formation control for unknown unmanned underwater vehicles (UUVs) under limited training data scenarios. First a distributed observer-based reference generator is designed, which decouples the exploration process of ADP for UUV from influencing neighboring UUVs’ learning outcomes. Second, to overcome data scarcity, we develop a dual-loop learning architecture: An inner-loop leverages least-squares (LS) iterations to find the LS solution from partial observations, reducing dependency on sufficient training data; An outer-loop uses policy iteration to optimize the control policy using Bellman equation. Theoretical analysis proves that the proposed DADP achieves convergence under limited training data. Comparing simulation results are provided to verify the advantages of the proposed approach over traditional ADP methods.
KW - Distributed adaptive dynamic programing
KW - Limited training data,
KW - Time-varying formation control
KW - Unmanned underwater vehicle
UR - https://www.scopus.com/pages/publications/105043174125
U2 - 10.1016/j.jfranklin.2026.108820
DO - 10.1016/j.jfranklin.2026.108820
M3 - Article
AN - SCOPUS:105043174125
SN - 0016-0032
VL - 363
JO - Journal of the Franklin Institute
JF - Journal of the Franklin Institute
IS - 12
M1 - 108820
ER -