RoG-DAgger: Rollout-Guided Post-Training for End-to-End Driving

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决端到端驾驶系统在策略引发状态下的安全问题,本文提出RoG-DAgger方法,通过短时运动预测构建高质量专家示例并适时调整控制权,以改善模型性能。
📝 Abstract
Recent end-to-end driving systems demonstrate strong performance on closed-loop benchmarks, yet are still predominantly trained on fixed expert-collected data using open-loop imitation learning. This training-inference mismatch leaves the policy vulnerable in policy-induced states, where accumulated errors can lead to safety-critical failures. A promising post-training approach to overcome this issue is Dataset Aggregation (DAgger), which gathers expert demonstrations in policy-induced states and subsequently fine-tunes the policy on the resulting aggregated dataset. Existing driving DAgger pipelines, however, face three challenges: i) the expert is restricted to a limited trajectory-and-speed solution space, ii) takeover may occur too early or too late relative to impending failures, and iii) privileged expert decisions may rely on information unavailable to the student. To address this, we introduce RoG-DAgger, a post-training framework that uses short-horizon kinematic rollouts to construct high-quality expert demonstrations in safety-critical states. Specifically, RoG-DAgger expands the expert's trajectory-and-speed solution space and evaluates candidate plans through rollout to construct preventive supervision. Moreover, it uses rollout solvability to time the takeover near the estimated point of no return. Lastly, it aligns the expert's field of view with that of the student to provide student-compatible supervision. Across in-distribution (including long-horizon) and out-of-distribution evaluations, RoG-DAgger improves the end-to-end model SimLingo by 5.3 driving-score points and 6.2 percentage points in success rate on Bench2Drive, doubles its driving score from 22 to 44 on Longest6 v2, and improves out-of-distribution success rate from 55\% to 66\% on Fail2Drive.
Problem

Research questions and friction points this paper is trying to address.

end-to-end driving
policy-induced states
safety-critical failures
dataset aggregation
imitation learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rollout-Guided Post-Training
Kinematic Rollouts
Preventive Supervision
Takeover Timing
Field of View Alignment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Liangyu Zhong
CARIAD SE, Volkswagen Group
J
Joachim Sicking
CARIAD SE, Volkswagen Group
F
Fabian Hueger
CARIAD SE, Volkswagen Group
Hanno Gottschalk
Hanno Gottschalk
Professor for Mathematical Modeling of Industrial Life Cycles, TU Berlin
Applied MathematicsComputer VisionMachine LearningMathematical Physics