Abstract
Urban road networks frequently experience recurrent congestion and non-recurrent disruptions, where individually optimal routing can degrade system-level efficiency under mixed compliance. Naive shortest-path guidance can be overly myopic in coupled traffic systems. As congestion is shaped by collective route choices, it may over-react to momentary link costs and trigger herding-like oscillations. This motivates routing controllers with decentralized execution that scale to many vehicles while better accounting for congestion externalities. We propose a routing framework centered on Traffic Cost Fields (TCF), which maintain factorized edge-wise signals as an interpretable interface between traffic states and routing decisions. A shared-parameter multi-agent actor-critic policy is trained with Proximal Policy Optimization (PPO) to make periodic intersection-level next-hop decisions from masked local observations, enabling decentralized execution at scale. Experiments on an abstracted road graph extracted from Shenzhen show improved average and tail travel-time performance over shortest-path baselines across multiple demand levels.
| Original language | English |
|---|---|
| Title of host publication | 2026 12th International Conference on Control, Automation and Robotics, ICCAR 2026 |
| Publisher | Institute of Electrical and Electronics Engineers Inc. |
| Pages | 1-7 |
| Number of pages | 7 |
| Edition | 2026 |
| ISBN (Electronic) | 9798319529350 |
| DOIs | |
| Publication status | Published - 2026 |
| Event | 2026 12th International Conference on Control, Automation and Robotics, ICCAR 2026 - Nagoya, Japan Duration: 8 Apr 2026 → 10 Apr 2026 |
Conference
| Conference | 2026 12th International Conference on Control, Automation and Robotics, ICCAR 2026 |
|---|---|
| Country/Territory | Japan |
| City | Nagoya |
| Period | 8/04/26 → 10/04/26 |
Keywords
- Path planning
- Reinforcement learning
- Routing control
- Transportation
Fingerprint
Dive into the research topics of 'Resilient Multi-Agent Deep Reinforcement Learning for Routing Control via Traffic Cost Fields'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver