TY - JOUR
T1 - Vision–Language Model-Enabled Dual-System for Autonomous Driving in Safety-Critical Transportation
AU - Li, Lin
AU - Fang, Jianwu
AU - Xue, Jianru
AU - Lv, Chen
N1 - Publisher Copyright:
© 2000-2011 IEEE.
PY - 2026
Y1 - 2026
N2 - Recent advances in vision–language models (VLMs) have enabled autonomous driving systems to perform high-level scene understanding and causal reasoning. However, existing VLM-based frameworks typically operate in an open-loop manner, lack real-time responsiveness, and exhibit weak coupling between language-grounded reasoning and continuous control, which limits their deployment in safety-critical environments. We introduce DualDrive, a novel dual-system vision–language architecture that fundamentally redefines the interaction between cognitive reasoning and driving control. Unlike prior dual-process or VLM-assisted approaches, DualDrive features two key innovations: 1) a semantic intent representation generated by a VLM-based reasoning module, enabling interpretable, context-aware decision guidance; and 2) a reasoning-to-control alignment mechanism, achieved through a dedicated offline reinforcement learning stage, which shapes the reasoning latent space to be directly executable by a lightweight policy network. This design ensures that high-level semantics produced by System 2 translate into stable, real-time control actions from System 1—without relying on multi-step buffering or asynchronous updates. Built upon a two-stage learning paradigm—driving knowledge pretraining and offline RL co-training—DualDrive establishes a tightly coupled cognitive-control pipeline that enhances both interpretability and robustness. Evaluations on the Bench2Drive benchmark show that DualDrive achieves state-of-the-art closed-loop performance, significantly outperforming existing methods in driving score, success rate, and safety-critical behavior, while maintaining real-time control efficiency. The results demonstrate that aligning VLM reasoning with policy execution is a promising direction for building trustworthy, human-like autonomous driving systems.
AB - Recent advances in vision–language models (VLMs) have enabled autonomous driving systems to perform high-level scene understanding and causal reasoning. However, existing VLM-based frameworks typically operate in an open-loop manner, lack real-time responsiveness, and exhibit weak coupling between language-grounded reasoning and continuous control, which limits their deployment in safety-critical environments. We introduce DualDrive, a novel dual-system vision–language architecture that fundamentally redefines the interaction between cognitive reasoning and driving control. Unlike prior dual-process or VLM-assisted approaches, DualDrive features two key innovations: 1) a semantic intent representation generated by a VLM-based reasoning module, enabling interpretable, context-aware decision guidance; and 2) a reasoning-to-control alignment mechanism, achieved through a dedicated offline reinforcement learning stage, which shapes the reasoning latent space to be directly executable by a lightweight policy network. This design ensures that high-level semantics produced by System 2 translate into stable, real-time control actions from System 1—without relying on multi-step buffering or asynchronous updates. Built upon a two-stage learning paradigm—driving knowledge pretraining and offline RL co-training—DualDrive establishes a tightly coupled cognitive-control pipeline that enhances both interpretability and robustness. Evaluations on the Bench2Drive benchmark show that DualDrive achieves state-of-the-art closed-loop performance, significantly outperforming existing methods in driving score, success rate, and safety-critical behavior, while maintaining real-time control efficiency. The results demonstrate that aligning VLM reasoning with policy execution is a promising direction for building trustworthy, human-like autonomous driving systems.
KW - Autonomous driving
KW - foundation model
KW - intelligent vehicles
KW - vision language model
UR - https://www.scopus.com/pages/publications/105044559718
U2 - 10.1109/TITS.2026.3708214
DO - 10.1109/TITS.2026.3708214
M3 - 文章
AN - SCOPUS:105044559718
SN - 1524-9050
JO - IEEE Transactions on Intelligent Transportation Systems
JF - IEEE Transactions on Intelligent Transportation Systems
ER -