A Penetration Method for UAV Based on Distributed Reinforcement Learning and Demonstrations

Kexv Li; Yue Wang; Xing Zhuang; Hao Yin; Xinyu Liu; Hanyu Li

doi:10.3390/drones7040232

A Penetration Method for UAV Based on Distributed Reinforcement Learning and Demonstrations

Kexv Li, Yue Wang^*, Xing Zhuang, Hao Yin, Xinyu Liu, Hanyu Li

^*此作品的通讯作者

机电学院

Beijing Institute of Technology

科研成果: 期刊稿件 › 文章 › 同行评审

2 引用（Scopus）

摘要

The penetration of unmanned aerial vehicles (UAVs) is an essential and important link in modern warfare. Enhancing UAV’s ability of autonomous penetration through machine learning has become a research hotspot. However, the current generation of autonomous penetration strategies for UAVs faces the problem of excessive sample demand. To reduce the sample demand, this paper proposes a combination policy learning (CPL) algorithm that combines distributed reinforcement learning and demonstrations. Innovatively, the action of the CPL algorithm is jointly determined by the initial policy obtained from demonstrations and the target policy in the asynchronous advantage actor-critic network, thus retaining the guiding role of demonstrations in the initial training. In a complex and unknown dynamic environment, 1000 training experiments and 500 test experiments were conducted for the CPL algorithm and related baseline algorithms. The results show that the CPL algorithm has the smallest sample demand, the highest convergence efficiency, and the highest success rate of penetration among all the algorithms, and has strong robustness in dynamic environments.

源语言	英语
文章编号	232
期刊	Drones
卷	7
期	4
DOI	https://doi.org/10.3390/drones7040232
出版状态	已出版 - 4月 2023

访问文件

10.3390/drones7040232

其它文件与链接

链接到 Scopus 的出版物

引用此

@article{ebb2092a48144da987a4f68aba722525,

title = "A Penetration Method for UAV Based on Distributed Reinforcement Learning and Demonstrations",

abstract = "The penetration of unmanned aerial vehicles (UAVs) is an essential and important link in modern warfare. Enhancing UAV{\textquoteright}s ability of autonomous penetration through machine learning has become a research hotspot. However, the current generation of autonomous penetration strategies for UAVs faces the problem of excessive sample demand. To reduce the sample demand, this paper proposes a combination policy learning (CPL) algorithm that combines distributed reinforcement learning and demonstrations. Innovatively, the action of the CPL algorithm is jointly determined by the initial policy obtained from demonstrations and the target policy in the asynchronous advantage actor-critic network, thus retaining the guiding role of demonstrations in the initial training. In a complex and unknown dynamic environment, 1000 training experiments and 500 test experiments were conducted for the CPL algorithm and related baseline algorithms. The results show that the CPL algorithm has the smallest sample demand, the highest convergence efficiency, and the highest success rate of penetration among all the algorithms, and has strong robustness in dynamic environments.",

keywords = "UAV penetration, asynchronous advantage actor-critic, demonstrations, distributed reinforcement learning",

author = "Kexv Li and Yue Wang and Xing Zhuang and Hao Yin and Xinyu Liu and Hanyu Li",

note = "Publisher Copyright: {\textcopyright} 2023 by the authors.",

year = "2023",

month = apr,

doi = "10.3390/drones7040232",

language = "English",

volume = "7",

journal = "Drones",

issn = "2504-446X",

publisher = "Multidisciplinary Digital Publishing Institute (MDPI)",

number = "4",

}

TY - JOUR

T1 - A Penetration Method for UAV Based on Distributed Reinforcement Learning and Demonstrations

AU - Li, Kexv

AU - Wang, Yue

AU - Zhuang, Xing

AU - Yin, Hao

AU - Liu, Xinyu

AU - Li, Hanyu

PY - 2023/4

Y1 - 2023/4

N2 - The penetration of unmanned aerial vehicles (UAVs) is an essential and important link in modern warfare. Enhancing UAV’s ability of autonomous penetration through machine learning has become a research hotspot. However, the current generation of autonomous penetration strategies for UAVs faces the problem of excessive sample demand. To reduce the sample demand, this paper proposes a combination policy learning (CPL) algorithm that combines distributed reinforcement learning and demonstrations. Innovatively, the action of the CPL algorithm is jointly determined by the initial policy obtained from demonstrations and the target policy in the asynchronous advantage actor-critic network, thus retaining the guiding role of demonstrations in the initial training. In a complex and unknown dynamic environment, 1000 training experiments and 500 test experiments were conducted for the CPL algorithm and related baseline algorithms. The results show that the CPL algorithm has the smallest sample demand, the highest convergence efficiency, and the highest success rate of penetration among all the algorithms, and has strong robustness in dynamic environments.

AB - The penetration of unmanned aerial vehicles (UAVs) is an essential and important link in modern warfare. Enhancing UAV’s ability of autonomous penetration through machine learning has become a research hotspot. However, the current generation of autonomous penetration strategies for UAVs faces the problem of excessive sample demand. To reduce the sample demand, this paper proposes a combination policy learning (CPL) algorithm that combines distributed reinforcement learning and demonstrations. Innovatively, the action of the CPL algorithm is jointly determined by the initial policy obtained from demonstrations and the target policy in the asynchronous advantage actor-critic network, thus retaining the guiding role of demonstrations in the initial training. In a complex and unknown dynamic environment, 1000 training experiments and 500 test experiments were conducted for the CPL algorithm and related baseline algorithms. The results show that the CPL algorithm has the smallest sample demand, the highest convergence efficiency, and the highest success rate of penetration among all the algorithms, and has strong robustness in dynamic environments.

KW - UAV penetration

KW - asynchronous advantage actor-critic

KW - demonstrations

KW - distributed reinforcement learning

UR - http://www.scopus.com/inward/record.url?scp=85153771926&partnerID=8YFLogxK

U2 - 10.3390/drones7040232

DO - 10.3390/drones7040232

M3 - Article

AN - SCOPUS:85153771926

SN - 2504-446X

VL - 7

JO - Drones

JF - Drones

IS - 4

M1 - 232

ER -

A Penetration Method for UAV Based on Distributed Reinforcement Learning and Demonstrations

摘要

访问文件

其它文件与链接

指纹

引用此