Skip to main navigation Skip to search Skip to main content

An Adaptive Multimodal End-to-End Autonomous Driving Framework Based on Behavior Discrepancy

  • Xiaokai Chen*
  • , Xiaoyu Wang
  • , Weihao Chen
  • , Jianping Gao
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • Henan University of Science and Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Vision-language models (VLMs) have recently advanced end-to-end autonomous driving by providing rich semantic understanding. However, existing approaches rarely consider practical constraints such as computational resource limits and inference latency, which are especially important in large-scale vehicle fleets or dense traffic scenarios with frequent VLM usage. To tackle this, we introduce a multimodal driving framework, G-VLM, which learns from the behavioral discrepancy between a lightweight controller and a powerful perception-language reasoning component. G-VLM incorporates a learned gating mechanism that adaptively determines when to activate the more expensive semantic processor, guided by scene context and confidence estimates. This enables a favorable trade-off between control accuracy and computational efficiency. Comprehensive evaluations on the CARLA closed-loop benchmark demonstrate that G-VLM achieves 96.8% of the fully-invoked VLM baseline’s driving score at only 12.2% of its computational cost, while outperforming compute-matched alternatives by up to 9.8 driving score points. The results highlight a promising direction where behavior-discrepancy-driven gating facilitates scalable and resource-aware autonomous driving, paving the way for efficient cloud-edge collaboration and intelligent mobility in smart city environments.

Original languageEnglish
JournalIEEE Transactions on Intelligent Transportation Systems
DOIs
Publication statusAccepted/In press - 2026
Externally publishedYes

Keywords

  • artificial intelligence
  • Autonomous driving
  • behavior discrepancy
  • gating mechanism
  • vision-language models

Fingerprint

Dive into the research topics of 'An Adaptive Multimodal End-to-End Autonomous Driving Framework Based on Behavior Discrepancy'. Together they form a unique fingerprint.

Cite this