Abstract
Recent advances in vision–language models (VLMs) have enabled autonomous driving systems to perform high-level scene understanding and causal reasoning. However, existing VLM-based frameworks typically operate in an open-loop manner, lack real-time responsiveness, and exhibit weak coupling between language-grounded reasoning and continuous control, which limits their deployment in safety-critical environments. We introduce DualDrive, a novel dual-system vision–language architecture that fundamentally redefines the interaction between cognitive reasoning and driving control. Unlike prior dual-process or VLM-assisted approaches, DualDrive features two key innovations: 1) a semantic intent representation generated by a VLM-based reasoning module, enabling interpretable, context-aware decision guidance; and 2) a reasoning-to-control alignment mechanism, achieved through a dedicated offline reinforcement learning stage, which shapes the reasoning latent space to be directly executable by a lightweight policy network. This design ensures that high-level semantics produced by System 2 translate into stable, real-time control actions from System 1—without relying on multi-step buffering or asynchronous updates. Built upon a two-stage learning paradigm—driving knowledge pretraining and offline RL co-training—DualDrive establishes a tightly coupled cognitive-control pipeline that enhances both interpretability and robustness. Evaluations on the Bench2Drive benchmark show that DualDrive achieves state-of-the-art closed-loop performance, significantly outperforming existing methods in driving score, success rate, and safety-critical behavior, while maintaining real-time control efficiency. The results demonstrate that aligning VLM reasoning with policy execution is a promising direction for building trustworthy, human-like autonomous driving systems.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Intelligent Transportation Systems |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- Autonomous driving
- foundation model
- intelligent vehicles
- vision language model
Fingerprint
Dive into the research topics of 'Vision–Language Model-Enabled Dual-System for Autonomous Driving in Safety-Critical Transportation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver