SafeGen: Goal-Conditioned Video Diffusion of Safety-Critical Scenarios for VLM-Based Autonomous Driving

πŸ“… 2026-07-21
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge that existing simulation-based methods for generating safety-critical driving scenarios struggle to realistically reproduce high-risk interactions involving vulnerable road users due to the sim-to-real gap. To overcome this limitation, the authors propose a goal-conditioned diffusion framework that leverages catastrophic end states as strong supervisory signals. By integrating vision-language models to analyze normal driving contexts and infer interaction vulnerabilities, the method enables context-anchored end-state reasoning and end-state-conditioned video evolution, yielding temporally coherent and physically plausible high-risk scenarios. Evaluated across three VLMAD systems, the approach improves the Judge Overall Score by an average of 24.25% and enhances the fine-tuned models’ performance on real-world scenarios by 15.9% on average.
πŸ“ Abstract
VLMs are increasingly deployed in AD systems, creating an urgent need for rigorous safety evaluation under rare yet safety-critical scenarios. Among these, interactions with vulnerable road users represent a major source of real-world failures. However, existing safety-critical scenario generation methods predominantly rely on simulator-based pipelines, which suffer from a substantial sim-to-real gap and often fail to capture realistic, diverse, and unforeseen human-vehicle interaction dynamics. We present SafeGen, a goal-conditioned diffusion framework for safety-critical scenario generation in VLMADs. Our key insight is to formulate scenario generation as a goal-conditioned diffusion process, where a predefined catastrophic end-state serves as a strong supervisory signal, guiding the generation of temporally coherent video trajectories that naturally evolve toward safety-critical outcomes. Building on this formulation, we introduce Context Grounded End State Reasoning, which leverages VLMs to analyze benign driving contexts and infer latent vulnerabilities in human-vehicle interactions, producing structured end-state specifications that induce high-risk scenarios. Conditioned on these targets, we further propose End State Conditioned Video Evolution, which grounds semantic threats into physically plausible visual dynamics. Specifically, we instantiate high-risk agents within the scene via depth-aware geometric projection, followed by boundary-conditioned diffusion to generate intermediate frames with consistent motion patterns and temporal coherence. Extensive experiments across 3 VLMADs demonstrate that SafeGen increases the Judge Overall Score, a metric using a VLM judge to evaluate VLMADs' understanding and decision-making, by 24.25% on average compared to SoTA baselines. Furthermore, fine-tuning a VLMAD improves performance in real-world driving scenes by an average of 15.9%.
Problem

Research questions and friction points this paper is trying to address.

safety-critical scenarios
VLM-based autonomous driving
sim-to-real gap
vulnerable road users
human-vehicle interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

goal-conditioned diffusion
safety-critical scenario generation
vision-language models
end-state reasoning
video diffusion
πŸ”Ž Similar Papers
No similar papers found.
J
Jiangfan Liu
Beihang University, China
Z
Zexuan Cui
Beihang University, China
Tianyuan Zhang
Tianyuan Zhang
MIT
Computer VisionMachine Learning
Zonglei Jing
Zonglei Jing
Beihang University
Machine LearningReinforcement LearningOptimal Control
Zonghao Ying
Zonghao Ying
SKLCCSE, BUAA
Trustworthy AI
Y
Yaoyuan Zhang
Beihang University, China
Jiakai Wang
Jiakai Wang
Zhongguancun Laboratory
Adversarial examplesTrustworthy AI
X
Xiaoqi Jiang
Chery Automobile Co., Ltd., China
A
Aishan Liu
Beihang University, China
X
Xianglong Liu
Beihang University, China; Zhongguancun Laboratory, China