TY - GEN
T1 - An improved PPO-GTrXL algorithm for missile evasion in partially observable air combat
AU - Gu, Yanhang
AU - Cui, Fang
AU - Guo, Xiang
AU - Wang, Bo
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Missile evasion is a critical aspect of aircraft survivability in modern air combat, particularly under the threat of agile, high-speed missiles with advanced guidance systems. However, conventional rule-based or optimal control approaches often fail to generalize in dynamic and partially observable environments. To address this challenge, we propose a novel missile evasion framework that integrates Proximal Policy Optimization (PPO) with Gated Transformer-XL (GTrXL). A high-fidelity simulation environment is developed, incorporating a realistic aircraft dynamics model and a proportional navigation-guided missile. We design a customized reward function to encourage survival-oriented behavior while adhering to flight constraints. The proposed PPO-GTrXL agent demonstrates enhanced adaptability and decision-making under uncertainty. Experimental results show that PPO-GTrXL improves the average return from -560.18 to 407.05 and raises the win rate from 63.3% to 85.7%, while also reducing performance variance compared to baseline PPO. These findings highlight the robustness and effectiveness of our method in complex aerial engagement scenarios.
AB - Missile evasion is a critical aspect of aircraft survivability in modern air combat, particularly under the threat of agile, high-speed missiles with advanced guidance systems. However, conventional rule-based or optimal control approaches often fail to generalize in dynamic and partially observable environments. To address this challenge, we propose a novel missile evasion framework that integrates Proximal Policy Optimization (PPO) with Gated Transformer-XL (GTrXL). A high-fidelity simulation environment is developed, incorporating a realistic aircraft dynamics model and a proportional navigation-guided missile. We design a customized reward function to encourage survival-oriented behavior while adhering to flight constraints. The proposed PPO-GTrXL agent demonstrates enhanced adaptability and decision-making under uncertainty. Experimental results show that PPO-GTrXL improves the average return from -560.18 to 407.05 and raises the win rate from 63.3% to 85.7%, while also reducing performance variance compared to baseline PPO. These findings highlight the robustness and effectiveness of our method in complex aerial engagement scenarios.
KW - Gated Transformer-XL
KW - missile evasion
KW - partial observability
KW - Proximal Policy Optimization
UR - https://www.scopus.com/pages/publications/105041058701
U2 - 10.1109/CAC67268.2025.11487564
DO - 10.1109/CAC67268.2025.11487564
M3 - Conference contribution
AN - SCOPUS:105041058701
T3 - Proceedings - 2025 China Automation Congress, CAC 2025
SP - 1482
EP - 1486
BT - Proceedings - 2025 China Automation Congress, CAC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 China Automation Congress, CAC 2025
Y2 - 26 September 2025 through 28 September 2025
ER -