Abstract
Vision-language models (VLMs) have recently advanced end-to-end autonomous driving by providing rich semantic understanding. However, existing approaches rarely consider practical constraints such as computational resource limits and inference latency, which are especially important in large-scale vehicle fleets or dense traffic scenarios with frequent VLM usage. To tackle this, we introduce a multimodal driving framework, G-VLM, which learns from the behavioral discrepancy between a lightweight controller and a powerful perception-language reasoning component. G-VLM incorporates a learned gating mechanism that adaptively determines when to activate the more expensive semantic processor, guided by scene context and confidence estimates. This enables a favorable trade-off between control accuracy and computational efficiency. Comprehensive evaluations on the CARLA closed-loop benchmark demonstrate that G-VLM achieves 96.8% of the fully-invoked VLM baseline’s driving score at only 12.2% of its computational cost, while outperforming compute-matched alternatives by up to 9.8 driving score points. The results highlight a promising direction where behavior-discrepancy-driven gating facilitates scalable and resource-aware autonomous driving, paving the way for efficient cloud-edge collaboration and intelligent mobility in smart city environments.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Intelligent Transportation Systems |
| DOIs | |
| Publication status | Accepted/In press - 2026 |
| Externally published | Yes |
Keywords
- artificial intelligence
- Autonomous driving
- behavior discrepancy
- gating mechanism
- vision-language models
Fingerprint
Dive into the research topics of 'An Adaptive Multimodal End-to-End Autonomous Driving Framework Based on Behavior Discrepancy'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver