Skip to main navigation Skip to search Skip to main content

Multi-resource Attack-Defense Strategy Optimization in the Blotto Game Based on Pool-PPO

  • Luying Chen
  • , Jie Hou
  • , Xianlin Zeng*
  • *Corresponding author for this work
  • Beijing Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Effective command of heterogeneous combat systems, comprising attack, reconnaissance, and electronic-warfare units, is critical to determining the outcome of modern conflicts, the core of which can be abstracted as a complex dynamic resource allocation problem. This paper models the problem as an extended, multi-stage Colonel Blotto game with multi-resources. To address the high-dimensional state space and the difficulty of solving for a mixed-strategy equilibrium, we propose Pool-PPO: an alternating self-play framework with a bounded historical opponent pool that mitigates non-stationarity and promotes mixed-strategy learning. Each outer iteration consists of two phases: (A) update the attacker against a defender snapshot sampled from the pool; (B) update the defender against the current attacker. The pool is maintained as a FIFO queue, periodically appending new defender snapshots and discarding the oldest. Experiments show that, relative to a symmetric PPO baseline, Pool-PPO yields a significantly higher expected payoff for the attacker while maintaining higher policy entropy, producing more randomized, less exploitable behavior that better approaches mixed-strategy Nash solutions. Overall, constructing and leveraging a historical opponent distribution within self-play offers an effective pathway to solving complex dynamic adversarial problems and obtaining robust, advantageous strategies.

Original languageEnglish
Title of host publicationProceedings of 2025 9th Chinese Conference on Swarm Intelligence and Cooperative Control - Swarm Optimization Technologies
EditorsYongzhao Hua, Yishi Liu, Rui Yan
PublisherSpringer Science and Business Media Deutschland GmbH
Pages399-411
Number of pages13
ISBN (Print)9789819583287
DOIs
Publication statusPublished - 2026
Externally publishedYes
Event9th Chinese Conference on Swarm Intelligence and Cooperative Control, CCSICC 2025 - Shanghai, China
Duration: 31 Oct 20253 Nov 2025

Publication series

NameLecture Notes in Electrical Engineering
Volume1606 LNEE
ISSN (Print)1876-1100
ISSN (Electronic)1876-1119

Conference

Conference9th Chinese Conference on Swarm Intelligence and Cooperative Control, CCSICC 2025
Country/TerritoryChina
CityShanghai
Period31/10/253/11/25

Keywords

  • Colonel Blotto Game
  • Deep Reinforcement Learning
  • Multi-resources
  • Multi-stage Dynamic Game

Fingerprint

Dive into the research topics of 'Multi-resource Attack-Defense Strategy Optimization in the Blotto Game Based on Pool-PPO'. Together they form a unique fingerprint.

Cite this