TY - GEN
T1 - A Visual Attention-Informed Scene Graph Approach for Predicting Risk of Vulnerable Road Users
AU - Fan, Sizhe
AU - Tao, Gang
AU - Lin, Yunlong
AU - Song, Ze
AU - Lu, Chao
AU - Gong, Jianwei
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Accurately predicting the risk of vulnerable road users (VRUs) in advance is critical for enhancing the safety of advanced driver assistance systems. However, most existing approaches rely solely on sensor data and fail to fully leverage driver attention, which is implicitly embedded in their visual behavior. In this paper, a novel multi-object risk prediction approach is proposed to incorporate driver gaze information into a spatio-temporal scene graph convolutional neural network for predicting the risk of VRUs. First, the attention state of each road user is calculated using driver gaze information. Then, attention features are embedded into a data-driven scene graph construction module. Finally, the resulting spatio-temporal scene graphs are processed using graph and temporal convolutions to predict the time to collision for each road user. Experimental results show that, compared with baseline methods, our approach maintains high prediction accuracy even over extended prediction horizons, enabling earlier and more reliable identification of high-risk VRUs in driving environments.
AB - Accurately predicting the risk of vulnerable road users (VRUs) in advance is critical for enhancing the safety of advanced driver assistance systems. However, most existing approaches rely solely on sensor data and fail to fully leverage driver attention, which is implicitly embedded in their visual behavior. In this paper, a novel multi-object risk prediction approach is proposed to incorporate driver gaze information into a spatio-temporal scene graph convolutional neural network for predicting the risk of VRUs. First, the attention state of each road user is calculated using driver gaze information. Then, attention features are embedded into a data-driven scene graph construction module. Finally, the resulting spatio-temporal scene graphs are processed using graph and temporal convolutions to predict the time to collision for each road user. Experimental results show that, compared with baseline methods, our approach maintains high prediction accuracy even over extended prediction horizons, enabling earlier and more reliable identification of high-risk VRUs in driving environments.
KW - gaze information
KW - risk prediction
KW - scene graph
KW - Vulnerable road users
UR - https://www.scopus.com/pages/publications/105041132913
U2 - 10.1109/CAC67268.2025.11487594
DO - 10.1109/CAC67268.2025.11487594
M3 - Conference contribution
AN - SCOPUS:105041132913
T3 - Proceedings - 2025 China Automation Congress, CAC 2025
SP - 2929
EP - 2934
BT - Proceedings - 2025 China Automation Congress, CAC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 China Automation Congress, CAC 2025
Y2 - 26 September 2025 through 28 September 2025
ER -