Planning in the Backbone: DiffAdapterVLA for Native Continuous Trajectory Generation with Driving VLMs

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出DiffAdapterVLA方法,通过在预训练驾驶视觉-语言模型的深层中注入轨迹标记,实现连续轨迹规划,提高规划效率与质量。
📝 Abstract
Pretrained driving vision-language models (VLMs) integrate visual, route, language, and driving context into rich driving priors, yet their representation objectives remain separated from continuous driving planning. Existing methods typically begin trajectory generation only after the VLM has formed a final condition, leaving depth-wise condition computation outside the stepwise formation of trajectory state. We introduce DiffAdapterVLA, which realizes Planning in the Backbone: it injects explicit trajectory tokens into selected VLM late layers, bringing trajectory state into backbone forward computation, where it co-evolves with driving conditions at different depths. Lightweight layer-wise DiffAdapters organize this computation into recursive trajectory refinement, while asymmetric joint attention preserves directed guidance from the condition stream to trajectory planning. By placing planning within existing backbone computation rather than relying on an independent trajectory planner, DiffAdapterVLA adapts only lightweight trajectory modules to turn existing driving priors into efficient continuous planning capability. NAVSIM results show that it achieves high-quality closed-loop planning with low end-to-end latency using few trainable parameters, and demonstrate that jointly evolving trajectory state and depth-wise driving conditions in VLM late-layer computation effectively realizes continuous trajectory planning.
Problem

Research questions and friction points this paper is trying to address.

driving vision-language models
continuous trajectory planning
trajectory generation
backbone computation
driving conditions
Innovation

Methods, ideas, or system contributions that make the work stand out.

DiffAdapterVLA
Planning in the Backbone
trajectory tokens
asymmetric joint attention
lightweight DiffAdapters
🔎 Similar Papers
2024-09-12IEEE Transactions on Automation Science and EngineeringCitations: 2
C
Changxin Lu
School of Remote Sensing and Information Engineering, Wuhan University
X
Xiaoliang Meng
School of Remote Sensing and Information Engineering, Wuhan University
Yu Wu
Yu Wu
University of Cambridge
machine learninghealth sensingmobile health
R
Rui Huang
Dongfeng Research & Development Institute
Honglin Li
Honglin Li
Westlake University
Computer VisionMultimodal LLMbiomedical image analysis
T
Tao Chen
Dongfeng Research & Development Institute
K
Kaixuan Zhou
Dongfeng Research & Development Institute
Y
Yadong Shao
Dongfeng Research & Development Institute