Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Online reinforcement learning struggles to adequately explore safety-critical scenarios in autonomous driving, often resulting in policies with insufficient robustness. This work proposes TPSP, a novel approach that, for the first time, integrates policy-aware scene encoding with threat-guided optimization to generate high-value training scenarios. By explicitly modeling policy–environment interactions, TPSP selectively perturbs critical objects in a manner aligned with the current policy’s weaknesses, enabling targeted and efficient safety-focused exploration. Evaluated on NAVSIM v2 using approximately 4 million kilometers of simulated driving data, TPSP demonstrates substantially improved safety performance compared to random or policy-agnostic perturbation strategies, significantly enhancing the efficiency of safe policy learning.
📝 Abstract
Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical driving scenes. The long-tailed nature of real-world traffic situations makes dangerous and rare interactions difficult to encounter through conventional sampling, limiting the ability of RL policies to learn robust safety behaviors. Existing methods improve training diversity by synthesizing challenging scenes or adversarial situations. However, these approaches typically optimize scene generation objectives separately from the evolving policy, without explicitly modeling how generated perturbations relate to the current policy's weaknesses and learning needs. In this paper, we propose Threat-guided Policy-aware Scene Perturbation (TPSP) for safe autonomous driving with online RL. TPSP introduces a policy-aware scene encoder to capture the interaction between policy behaviors and surrounding environments, enabling scene perturbation aligned with the current policy. Based on this representation, TPSP selectively perturbs critical objects rather than applying uniform modifications across the scene. Furthermore, we develop a threat-guided optimization strategy that evaluates perturbed scenes through threat-level differences between policy rollouts on original and perturbed scenes, guiding the generation of safety-critical scenes with higher training value. Comprehensive experiments demonstrate that TPSP improves safety learning efficiency, achieving strong safety performance on NAVSIM v2 with approximately 4 million kilometers of simulated driving data. Ablation studies verify that policy-aware targeted perturbations provide more informative safety-critical experiences than random or policy-unaware strategies, enabling safer driving under limited interaction budgets.
Problem

Research questions and friction points this paper is trying to address.

safe autonomous driving
online reinforcement learning
safety-critical scenes
long-tailed traffic scenarios
policy safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

policy-aware perturbation
threat-guided optimization
scene synthesis
safe reinforcement learning
autonomous driving
💼 Related Jobs
No related jobs found.
X
Xincong Hu
Nanjing University, Nanjing, Jiangsu, China
L
Lei Ou
Nanjing University, Nanjing, Jiangsu, China
M
Maosen Li
Yinwang Intelligent Technology Co., Ltd., China
J
Jingtao Zhang
Yinwang Intelligent Technology Co., Ltd., China
L
Liguo Hou
Yinwang Intelligent Technology Co., Ltd., China
Zongzhang Zhang
Zongzhang Zhang
Nanjing University
Artificial IntelligenceReinforcement LearningProbabilistic PlanningMulti-Agent Systems