Qwen-Drive-1.0:迈向自动驾驶视觉语言基础模型的初步探索
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
August 31, 2026
作者: Xin Zhou, Zongchuang Zhao, Zhibo Yang, Mingsheng Li, Humen Zhong, Shuai Bai, Du Chu, Ruizhe Chen, Zhaohai Li, Jun Tang, Qiuyue Wang, Mingkun Yang, Jiazhao Zhang, Dayiheng Liu, Dingkang Liang, Xiang Bai
cs.AI
摘要
我们提出了 Qwen-Drive-1.0,这是迈向自动驾驶视觉语言基础模型的初步探索。Qwen-Drive-1.0 保留了预训练视觉语言模型(VLM)的架构,并在统一框架内整合了三维感知、视觉问答与运动规划。一个外部的鸟瞰图(BEV)感知头联合执行三维目标检测、语义占用预测和 BEV 地图分割。它作为对共享表示中可用三维信息的探针,并为三维场景结构提供显式且可检查的接口。一个规划专家(Planning Expert)以共享的 VLM 表示为条件,生成未来的自车轨迹。分阶段的训练策略将驾驶监督与通用视觉语言数据相结合,以获取驾驶特定能力,同时帮助保持广泛的视觉理解与指令跟随能力。实验表明,该方法展现出强大的三维感知与驾驶场景理解能力,同时在很大程度上保持了通用视觉语言能力。在开环、伪闭环和闭环设置下的综合评估进一步显示,其运动规划性能具有很强的竞争力。
English
We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-eye-view (BEV) perception head jointly performs 3D object detection, semantic occupancy prediction, and BEV map segmentation. It serves as a probe of the 3D information accessible from the shared representations and provides an explicit, inspectable interface to 3D scene structure. A Planning Expert conditions on shared VLM representations to generate future ego trajectories. A staged training recipe combines driving supervision with general-purpose vision-language data to acquire driving-specific competence while helping preserve broad visual understanding and instruction-following capabilities. Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability. Comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.