Qwen-Drive-1.0: 자율주행을 위한 비전-언어 파운데이션 모델의 첫걸음

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

August 31, 2026
저자: Xin Zhou, Zongchuang Zhao, Zhibo Yang, Mingsheng Li, Humen Zhong, Shuai Bai, Du Chu, Ruizhe Chen, Zhaohai Li, Jun Tang, Qiuyue Wang, Mingkun Yang, Jiazhao Zhang, Dayiheng Liu, Dingkang Liang, Xiang Bai
cs.AI

초록

본 논문에서는 자율주행을 위한 비전-언어 파운데이션 모델 개발의 첫 단계로서 Qwen-Drive-1.0을 제안한다. Qwen-Drive-1.0은 사전 학습된 비전-언어 모델(VLM)의 아키텍처를 유지하면서 3D 인식, 시각 질의응답, 모션 플래닝을 하나의 통합 프레임워크 안에 결합한다. 외부의 조감도(BEV) 인식 헤드는 3D 객체 탐지, 의미론적 점유 예측, BEV 맵 분할을 동시에 수행한다. 이 헤드는 공유된 VLM 표현으로부터 접근 가능한 3D 정보를 점검하는 프로브 역할을 하며, 3D 장면 구조에 대한 명시적이고 검사 가능한 인터페이스를 제공한다. 플래닝 전문가(Planning Expert)는 공유 VLM 표현을 조건으로 삼아 미래의 자차 궤적을 생성한다. 단계적 훈련 방식은 주행 지도(supervision)와 범용 비전-언어 데이터를 결합하여 주행 특화 역량을 습득하는 동시에 폭넓은 시각 이해 및 지시 수행 능력의 보존을 돕는다. 실험 결과, 일반적인 비전-언어 능력이 대부분 유지되면서 강력한 3D 인식과 주행 장면 이해 성능이 입증되었다. 또한 개루프, 유사 폐루프, 폐루프 설정을 아우르는 종합 평가를 통해 매우 경쟁력 있는 모션 플래닝 성능을 확인하였다.
English
We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-eye-view (BEV) perception head jointly performs 3D object detection, semantic occupancy prediction, and BEV map segmentation. It serves as a probe of the 3D information accessible from the shared representations and provides an explicit, inspectable interface to 3D scene structure. A Planning Expert conditions on shared VLM representations to generate future ego trajectories. A staged training recipe combines driving supervision with general-purpose vision-language data to acquire driving-specific competence while helping preserve broad visual understanding and instruction-following capabilities. Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability. Comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.
PDF3373September 3, 2026