Qwen-Drive-1.0:朝向自動駕駛視覺語言基礎模型的初步探索

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

August 31, 2026
作者: Xin Zhou, Zongchuang Zhao, Zhibo Yang, Mingsheng Li, Humen Zhong, Shuai Bai, Du Chu, Ruizhe Chen, Zhaohai Li, Jun Tang, Qiuyue Wang, Mingkun Yang, Jiazhao Zhang, Dayiheng Liu, Dingkang Liang, Xiang Bai
cs.AI

摘要

我們提出 Qwen-Drive-1.0,這是朝向自動駕駛視覺-語言基礎模型的第一步。Qwen-Drive-1.0 保留了預訓練視覺-語言模型(VLM)的架構,並在統一框架中整合了 3D 感知、視覺問答與運動規劃。一個外部鳥瞰圖(BEV)感知頭聯合執行 3D 物體偵測、語義佔用預測與 BEV 地圖分割。它作為從共享表徵中可獲取 3D 資訊的探針,並提供一個明確且可檢視的介面來理解 3D 場景結構。一個規劃專家模組以共享的 VLM 表徵為條件,生成未來的自車軌跡。分階段的訓練配方將駕駛監督與通用視覺-語言資料結合,以獲得駕駛特定能力,同時有助於保留廣泛的視覺理解與指令跟隨能力。實驗表明,模型具備強大的 3D 感知與駕駛場景理解能力,同時在很大程度上保留了通用視覺-語言能力。在開環、偽閉環與閉環設定下的全面評估進一步顯示出其極具競爭力的運動規劃表現。
English
We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-eye-view (BEV) perception head jointly performs 3D object detection, semantic occupancy prediction, and BEV map segmentation. It serves as a probe of the 3D information accessible from the shared representations and provides an explicit, inspectable interface to 3D scene structure. A Planning Expert conditions on shared VLM representations to generate future ego trajectories. A staged training recipe combines driving supervision with general-purpose vision-language data to acquire driving-specific competence while helping preserve broad visual understanding and instruction-following capabilities. Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability. Comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.
PDF3373September 3, 2026