TY - GEN
T1 - Multi-resource Attack-Defense Strategy Optimization in the Blotto Game Based on Pool-PPO
AU - Chen, Luying
AU - Hou, Jie
AU - Zeng, Xianlin
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - Effective command of heterogeneous combat systems, comprising attack, reconnaissance, and electronic-warfare units, is critical to determining the outcome of modern conflicts, the core of which can be abstracted as a complex dynamic resource allocation problem. This paper models the problem as an extended, multi-stage Colonel Blotto game with multi-resources. To address the high-dimensional state space and the difficulty of solving for a mixed-strategy equilibrium, we propose Pool-PPO: an alternating self-play framework with a bounded historical opponent pool that mitigates non-stationarity and promotes mixed-strategy learning. Each outer iteration consists of two phases: (A) update the attacker against a defender snapshot sampled from the pool; (B) update the defender against the current attacker. The pool is maintained as a FIFO queue, periodically appending new defender snapshots and discarding the oldest. Experiments show that, relative to a symmetric PPO baseline, Pool-PPO yields a significantly higher expected payoff for the attacker while maintaining higher policy entropy, producing more randomized, less exploitable behavior that better approaches mixed-strategy Nash solutions. Overall, constructing and leveraging a historical opponent distribution within self-play offers an effective pathway to solving complex dynamic adversarial problems and obtaining robust, advantageous strategies.
AB - Effective command of heterogeneous combat systems, comprising attack, reconnaissance, and electronic-warfare units, is critical to determining the outcome of modern conflicts, the core of which can be abstracted as a complex dynamic resource allocation problem. This paper models the problem as an extended, multi-stage Colonel Blotto game with multi-resources. To address the high-dimensional state space and the difficulty of solving for a mixed-strategy equilibrium, we propose Pool-PPO: an alternating self-play framework with a bounded historical opponent pool that mitigates non-stationarity and promotes mixed-strategy learning. Each outer iteration consists of two phases: (A) update the attacker against a defender snapshot sampled from the pool; (B) update the defender against the current attacker. The pool is maintained as a FIFO queue, periodically appending new defender snapshots and discarding the oldest. Experiments show that, relative to a symmetric PPO baseline, Pool-PPO yields a significantly higher expected payoff for the attacker while maintaining higher policy entropy, producing more randomized, less exploitable behavior that better approaches mixed-strategy Nash solutions. Overall, constructing and leveraging a historical opponent distribution within self-play offers an effective pathway to solving complex dynamic adversarial problems and obtaining robust, advantageous strategies.
KW - Colonel Blotto Game
KW - Deep Reinforcement Learning
KW - Multi-resources
KW - Multi-stage Dynamic Game
UR - https://www.scopus.com/pages/publications/105040523556
U2 - 10.1007/978-981-95-8329-4_32
DO - 10.1007/978-981-95-8329-4_32
M3 - Conference contribution
AN - SCOPUS:105040523556
SN - 9789819583287
T3 - Lecture Notes in Electrical Engineering
SP - 399
EP - 411
BT - Proceedings of 2025 9th Chinese Conference on Swarm Intelligence and Cooperative Control - Swarm Optimization Technologies
A2 - Hua, Yongzhao
A2 - Liu, Yishi
A2 - Yan, Rui
PB - Springer Science and Business Media Deutschland GmbH
T2 - 9th Chinese Conference on Swarm Intelligence and Cooperative Control, CCSICC 2025
Y2 - 31 October 2025 through 3 November 2025
ER -