TY - GEN
T1 - A Cognition-Driven Network for Driver Gaze Prediction in Intelligent Vehicles
AU - Song, Ze
AU - Lu, Chao
AU - Wu, Tongshuai
AU - Fan, Sizhe
AU - Lin, Yunlong
AU - Gong, Jianwei
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Driver gaze prediction is a key research area in human-machine co-driving. It enables co-driving systems decide whether to take over control by predicting driver gaze and identifying possible driver distraction. However, most existing methods only focus on bottom-up attention, which solely relies on visual input from the driving scene, neglecting the top-down attention related to the driver cognition of the scene. In this paper, a novel dual-pipeline cognition-driven network named CoGNet is proposed to improve the accuracy of the gaze prediction. Driver scene cognition is derived from a top-down pipeline that takes as input semantic segmentation and internal vehicle information, such as the ego vehicle's speed and angular velocity. A cognitive-visual attention module is then introduced to fuse driver scene cognition with visual features extracted from the bottom-up pipeline. In addition, a comprehensive eye-tracking dataset is constructed with frame-level vehicle motion data such as speed and angular velocity, addressing the absence of such data in existing open-source datasets. Experiments on this dataset show that, CoGNet outperforms state-of-the-art methods across most metrics, resulting in more accurate driver gaze prediction.
AB - Driver gaze prediction is a key research area in human-machine co-driving. It enables co-driving systems decide whether to take over control by predicting driver gaze and identifying possible driver distraction. However, most existing methods only focus on bottom-up attention, which solely relies on visual input from the driving scene, neglecting the top-down attention related to the driver cognition of the scene. In this paper, a novel dual-pipeline cognition-driven network named CoGNet is proposed to improve the accuracy of the gaze prediction. Driver scene cognition is derived from a top-down pipeline that takes as input semantic segmentation and internal vehicle information, such as the ego vehicle's speed and angular velocity. A cognitive-visual attention module is then introduced to fuse driver scene cognition with visual features extracted from the bottom-up pipeline. In addition, a comprehensive eye-tracking dataset is constructed with frame-level vehicle motion data such as speed and angular velocity, addressing the absence of such data in existing open-source datasets. Experiments on this dataset show that, CoGNet outperforms state-of-the-art methods across most metrics, resulting in more accurate driver gaze prediction.
KW - Human-machine co-driving
KW - driver scene cognition
KW - gaze prediction
KW - top-down attention
UR - https://www.scopus.com/pages/publications/105041043448
U2 - 10.1109/CAC67268.2025.11487706
DO - 10.1109/CAC67268.2025.11487706
M3 - Conference contribution
AN - SCOPUS:105041043448
T3 - Proceedings - 2025 China Automation Congress, CAC 2025
SP - 6756
EP - 6761
BT - Proceedings - 2025 China Automation Congress, CAC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 China Automation Congress, CAC 2025
Y2 - 26 September 2025 through 28 September 2025
ER -