Abstract
Eco-driving at signalized intersections must balance energy efficiency, traffic efficiency, and safety under dynamic signal and traffic constraints. Although deep reinforcement learning (DRL) has shown promise for this task, policies trained in nominal environments may lose reliability when unseen signal timings or abrupt preceding-vehicle maneuvers expose experience-sparse boundary states. Under the conventional DRL train-and-test paradigm, such evaluation failures are usually recorded as terminal outcomes rather than reused for policy improvement. To address this missing data loop, this paper proposes a closed-loop offline policy improvement framework that recycles evaluation trajectories for successor-policy refinement. Specifically, safe and failure trajectories from a predecessor policy are collected into a mixed-quality offline dataset and reused through offline reinforcement learning, enabling value-guided improvement while regularizing policy updates toward data-supported actions. Results are evaluated under randomized signal-timing scenarios and an emergency-braking case. The proposed framework improves safety robustness over conventional DRL baselines and matches or exceeds the safety performance of a safety-constrained DRL baseline, with less conservative and less reactive behavior. The refined policy attains 6.65 ± 0.71 kWh/100 km, the lowest among all deployable methods, corresponding to 92.7% of the Dynamic Programming based theoretical reference computed under ideal information conditions. Vehicle-in-the-loop tests further indicate that smoother longitudinal behavior remains observable after real execution dynamics are introduced.
| Original language | English |
|---|---|
| Article number | 141837 |
| Journal | Energy |
| Volume | 360 |
| DOIs | |
| Publication status | Published - 30 Sept 2026 |
Keywords
- Closed-loop offline policy improvement
- Connected and automated vehicles
- Eco-driving
- Offline reinforcement learning
- Out-of-distribution generalization
Fingerprint
Dive into the research topics of 'Closed-loop offline policy improvement for eco-driving at signalized intersections via evaluation-trajectory recycling'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver