Qwen-Drive-1.0:自動運転のための視覚言語基盤モデルに向けた初期の一歩

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

August 31, 2026
著者: Xin Zhou, Zongchuang Zhao, Zhibo Yang, Mingsheng Li, Humen Zhong, Shuai Bai, Du Chu, Ruizhe Chen, Zhaohai Li, Jun Tang, Qiuyue Wang, Mingkun Yang, Jiazhao Zhang, Dayiheng Liu, Dingkang Liang, Xiang Bai
cs.AI

要旨

Here is the translation to Japanese: 本稿では、自動運転のための視覚言語基盤モデルへの初期段階として、Qwen-Drive-1.0を提案する。Qwen-Drive-1.0は、事前学習済み視覚言語モデル(VLM)のアーキテクチャを保持しつつ、3次元知覚、視覚質問応答、および運動計画を統合フレームワーク内に組み込む。外部のBird's-Eye-View(BEV)知覚ヘッドは、3次元物体検出、セマンティック占有予測、およびBEVマップセグメンテーションを併せて実行する。これは、共有表現からアクセス可能な3次元情報の探査手段として機能し、3次元シーン構造への明示的かつ検証可能なインターフェースを提供する。Plan Expertは共有VLM表現に基づいて将来の自車軌跡を生成する。段階的学習レシピは、運転タスクの教師信号を汎用の視覚言語データと組み合わせることで、運転特化の能力を獲得すると同時に、広範な視覚理解および命令追従能力の維持を支援する。実験により、汎用の視覚言語能力を概ね維持しつつ、強力な3次元知覚および運転シーン理解を実証する。オープンループ、疑似クローズドループ、およびクローズドループ設定にわたる包括的な評価により、非常に競争力のある運動計画性能がさらに示される。 Note: The text was segmented into logical paragraphs for readability while preserving the original structure and technical terminology. Key terms such as "VLM," "BEV," "3D," and model names (Qwen-Drive-1.0) remain in English per academic convention in Japanese literature.
English
We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-eye-view (BEV) perception head jointly performs 3D object detection, semantic occupancy prediction, and BEV map segmentation. It serves as a probe of the 3D information accessible from the shared representations and provides an explicit, inspectable interface to 3D scene structure. A Planning Expert conditions on shared VLM representations to generate future ego trajectories. A staged training recipe combines driving supervision with general-purpose vision-language data to acquire driving-specific competence while helping preserve broad visual understanding and instruction-following capabilities. Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability. Comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.
PDF3373September 3, 2026