Skip to main navigation Skip to search Skip to main content

Vision–Language Model-Enabled Dual-System for Autonomous Driving in Safety-Critical Transportation

  • Nanyang Technological University
  • Xi'an Jiaotong University

Research output: Contribution to journalArticlepeer-review

Abstract

Recent advances in vision–language models (VLMs) have enabled autonomous driving systems to perform high-level scene understanding and causal reasoning. However, existing VLM-based frameworks typically operate in an open-loop manner, lack real-time responsiveness, and exhibit weak coupling between language-grounded reasoning and continuous control, which limits their deployment in safety-critical environments. We introduce DualDrive, a novel dual-system vision–language architecture that fundamentally redefines the interaction between cognitive reasoning and driving control. Unlike prior dual-process or VLM-assisted approaches, DualDrive features two key innovations: 1) a semantic intent representation generated by a VLM-based reasoning module, enabling interpretable, context-aware decision guidance; and 2) a reasoning-to-control alignment mechanism, achieved through a dedicated offline reinforcement learning stage, which shapes the reasoning latent space to be directly executable by a lightweight policy network. This design ensures that high-level semantics produced by System 2 translate into stable, real-time control actions from System 1—without relying on multi-step buffering or asynchronous updates. Built upon a two-stage learning paradigm—driving knowledge pretraining and offline RL co-training—DualDrive establishes a tightly coupled cognitive-control pipeline that enhances both interpretability and robustness. Evaluations on the Bench2Drive benchmark show that DualDrive achieves state-of-the-art closed-loop performance, significantly outperforming existing methods in driving score, success rate, and safety-critical behavior, while maintaining real-time control efficiency. The results demonstrate that aligning VLM reasoning with policy execution is a promising direction for building trustworthy, human-like autonomous driving systems.

Original languageEnglish
JournalIEEE Transactions on Intelligent Transportation Systems
DOIs
StateAccepted/In press - 2026

Keywords

  • Autonomous driving
  • foundation model
  • intelligent vehicles
  • vision language model

Fingerprint

Dive into the research topics of 'Vision–Language Model-Enabled Dual-System for Autonomous Driving in Safety-Critical Transportation'. Together they form a unique fingerprint.

Cite this