Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出Qwen-Drive-1.0模型,整合3D感知、视觉问答与路径规划于统一框架,通过阶段性训练方法提升自动驾驶能力。
📝 Abstract
We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-eye-view (BEV) perception head jointly performs 3D object detection, semantic occupancy prediction, and BEV map segmentation. It serves as a probe of the 3D information accessible from the shared representations and provides an explicit, inspectable interface to 3D scene structure. A Planning Expert conditions on shared VLM representations to generate future ego trajectories. A staged training recipe combines driving supervision with general-purpose vision-language data to acquire driving-specific competence while helping preserve broad visual understanding and instruction-following capabilities. Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability. Comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.
Problem

Research questions and friction points this paper is trying to address.

vision-language model
autonomous driving
3D perception
visual question answering
motion planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language foundation model
unified framework
BEV perception head
staged training recipe
🔎 Similar Papers
No similar papers found.