TY - GEN
T1 - Learning to Model Diverse Interactive Traffic with Driving Tendency-Guided Policy Optimization
AU - Fan, Jialin
AU - Ni, Ying
AU - Yang, Yuhao
AU - Zheng, Wentao
AU - Sun, Jie
AU - Sun, Jian
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - The safe deployment of autonomous vehicles (AVs) into real-world traffic requires robust interaction with human drivers exhibiting heterogeneous behavioral tendencies, spanning from rational cooperation to adversarial aggression. Existing simulation frameworks often lack the capacity to systematically model such behavioral diversity, limiting their applicability for rigorous A V evaluation. To address this challenge, we propose a multi-agent reinforcement learning framework that generates dynamically controllable traffic through Tendency-Guided Policy Optimization (TGPO). Central to TGPO is the Adversary-Rationality-Tendency (ART), a continuous hyperparameter that enables fine-grained control over the spectrum of driving behaviors by fusing separately learned adversarial and rational value functions. Furthermore, we design an ART -guided policy network incorporating multi-head mechanisms to resolve high-dimensional multi-agent observations, adaptively prioritizing context features aligned with assigned driving tendencies. Extensive experiments across urban and highway scenarios demonstrate that TGPO generates traffic with enhanced behavioral controllability and diversity, which provides a scalable solution for simulating realistic interactions with various driving tendencies, thereby facilitating the development of AV systems capable of handling complex real-world corner cases.
AB - The safe deployment of autonomous vehicles (AVs) into real-world traffic requires robust interaction with human drivers exhibiting heterogeneous behavioral tendencies, spanning from rational cooperation to adversarial aggression. Existing simulation frameworks often lack the capacity to systematically model such behavioral diversity, limiting their applicability for rigorous A V evaluation. To address this challenge, we propose a multi-agent reinforcement learning framework that generates dynamically controllable traffic through Tendency-Guided Policy Optimization (TGPO). Central to TGPO is the Adversary-Rationality-Tendency (ART), a continuous hyperparameter that enables fine-grained control over the spectrum of driving behaviors by fusing separately learned adversarial and rational value functions. Furthermore, we design an ART -guided policy network incorporating multi-head mechanisms to resolve high-dimensional multi-agent observations, adaptively prioritizing context features aligned with assigned driving tendencies. Extensive experiments across urban and highway scenarios demonstrate that TGPO generates traffic with enhanced behavioral controllability and diversity, which provides a scalable solution for simulating realistic interactions with various driving tendencies, thereby facilitating the development of AV systems capable of handling complex real-world corner cases.
UR - https://www.scopus.com/pages/publications/105014240890
U2 - 10.1109/IV64158.2025.11097406
DO - 10.1109/IV64158.2025.11097406
M3 - Conference contribution
AN - SCOPUS:105014240890
T3 - IEEE Intelligent Vehicles Symposium, Proceedings
SP - 2047
EP - 2053
BT - IV 2025 - 36th IEEE Intelligent Vehicles Symposium
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 36th IEEE Intelligent Vehicles Symposium, IV 2025
Y2 - 22 June 2025 through 25 June 2025
ER -