RAF-VLA: Representation Alignment with the Future for End-to-End Autonomous Driving

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决VLA模型在自动驾驶中因未来场景生成导致的训练负担和推理延迟问题,提出RAF-VLA方法,通过直接对齐未来帧表示来优化内部表征。
📝 Abstract
Recent Vision-Language-Action (VLA) models for autonomous driving have incorporated world modeling by predicting future driving scenes alongside driving actions, demonstrating strong planning performance. Future driving scenes are utilized as dense supervision, encouraging the policy to learn rich internal representations useful for planning. However, these World-Modeling VLAs rely on explicit future generation to learn such representations, thereby introducing two key limitations: additional training burden and inference latency. To address these limitations, we propose RAF-VLA (Representation Alignment with the Future), a VLA-based autonomous driving framework that shapes planning-relevant internal representations through direct guidance from future-frame representations. RAF-VLA employs Future-Aligned Supervised Fine-Tuning, in which a straightforward regularization aligns the policy's hidden states with future-frame representations obtained from a pretrained world encoder while learning driving actions. This simple alignment allows RAF-VLA to avoid the training burden and inference latency associated with future generation. Extensive experiments on the NAVSIM benchmark show that RAF-VLA achieves competitive planning performance against state-of-the-art VLA planners with substantially fewer training samples seen. Moreover, RAF-VLA incurs only 3.8% training overhead and a negligible 1 ms inference overhead.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
autonomous driving
future generation
training burden
inference latency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Future-Aligned Supervised Fine-Tuning
representation alignment
autonomous driving
inference latency
training burden
🔎 Similar Papers
No similar papers found.
D
Dogun Kim
Graduate School of Mobility, Korea Advanced Institute of Science & Technology (KAIST), Daejeon, Korea
Y
Yongjae Lee
Graduate School of Mobility, Korea Advanced Institute of Science & Technology (KAIST), Daejeon, Korea
J
Joonhee Lim
Robotics Program, Korea Advanced Institute of Science & Technology (KAIST), Daejeon, Korea
Y
Yeina Lee
Robotics Program, Korea Advanced Institute of Science & Technology (KAIST), Daejeon, Korea
J
Junhyeok Park
Graduate School of Mobility, Korea Advanced Institute of Science & Technology (KAIST), Daejeon, Korea
M
Moogeun Park
Robotics Program, Korea Advanced Institute of Science & Technology (KAIST), Daejeon, Korea
Dongsuk Kum
Dongsuk Kum
KAIST
Vehicle Dynamics & Control