TY - JOUR
T1 - An Adaptive Multimodal End-to-End Autonomous Driving Framework Based on Behavior Discrepancy
AU - Chen, Xiaokai
AU - Wang, Xiaoyu
AU - Chen, Weihao
AU - Gao, Jianping
N1 - Publisher Copyright:
© 2000-2011 IEEE.
PY - 2026
Y1 - 2026
N2 - Vision-language models (VLMs) have recently advanced end-to-end autonomous driving by providing rich semantic understanding. However, existing approaches rarely consider practical constraints such as computational resource limits and inference latency, which are especially important in large-scale vehicle fleets or dense traffic scenarios with frequent VLM usage. To tackle this, we introduce a multimodal driving framework, G-VLM, which learns from the behavioral discrepancy between a lightweight controller and a powerful perception-language reasoning component. G-VLM incorporates a learned gating mechanism that adaptively determines when to activate the more expensive semantic processor, guided by scene context and confidence estimates. This enables a favorable trade-off between control accuracy and computational efficiency. Comprehensive evaluations on the CARLA closed-loop benchmark demonstrate that G-VLM achieves 96.8% of the fully-invoked VLM baseline’s driving score at only 12.2% of its computational cost, while outperforming compute-matched alternatives by up to 9.8 driving score points. The results highlight a promising direction where behavior-discrepancy-driven gating facilitates scalable and resource-aware autonomous driving, paving the way for efficient cloud-edge collaboration and intelligent mobility in smart city environments.
AB - Vision-language models (VLMs) have recently advanced end-to-end autonomous driving by providing rich semantic understanding. However, existing approaches rarely consider practical constraints such as computational resource limits and inference latency, which are especially important in large-scale vehicle fleets or dense traffic scenarios with frequent VLM usage. To tackle this, we introduce a multimodal driving framework, G-VLM, which learns from the behavioral discrepancy between a lightweight controller and a powerful perception-language reasoning component. G-VLM incorporates a learned gating mechanism that adaptively determines when to activate the more expensive semantic processor, guided by scene context and confidence estimates. This enables a favorable trade-off between control accuracy and computational efficiency. Comprehensive evaluations on the CARLA closed-loop benchmark demonstrate that G-VLM achieves 96.8% of the fully-invoked VLM baseline’s driving score at only 12.2% of its computational cost, while outperforming compute-matched alternatives by up to 9.8 driving score points. The results highlight a promising direction where behavior-discrepancy-driven gating facilitates scalable and resource-aware autonomous driving, paving the way for efficient cloud-edge collaboration and intelligent mobility in smart city environments.
KW - artificial intelligence
KW - Autonomous driving
KW - behavior discrepancy
KW - gating mechanism
KW - vision-language models
UR - https://www.scopus.com/pages/publications/105044063913
U2 - 10.1109/TITS.2026.3706626
DO - 10.1109/TITS.2026.3706626
M3 - Article
AN - SCOPUS:105044063913
SN - 1524-9050
JO - IEEE Transactions on Intelligent Transportation Systems
JF - IEEE Transactions on Intelligent Transportation Systems
ER -