Gradient-Free Neural Hamilton-Jacobi Reachability for Scalable Safety-Critical Control

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种无梯度神经Hamilton-Jacobi可达性框架,通过Bellman-Isaacs值传播学习安全关键控制的可达管,解决了高维非线性系统中的安全证书和鲁棒控制器合成问题。
📝 Abstract
Hamilton-Jacobi (HJ) reachability provides a principled framework for synthesizing safety certificates and robust controllers for safety-critical robotic systems. However, applying reachability analysis to high-dimensional nonlinear systems remains challenging: classical grid-based solvers suffer from the curse of dimensionality, continuous-time neural solvers require accurate spatial value gradients, and reinforcement-learning-based approaches often suffer from weak boundary anchoring and non-stationary adversarial policy optimization. We propose a discrete-time neural reachability framework for control-disturbance-affine systems that learns backward reachable tubes (BRTs) and backward reach-avoid tubes (BRATs) through Bellman-Isaacs value propagation. Our key idea is to combine equation-driven self-supervision with structured policy learning: rather than computing explicit PDE-gradients, we exploit the bang-bang structure of optimal safety interventions to construct approximate teacher actions from gradient-free value probes, converting adversarial actor learning into supervised policy learning. To stabilize long-horizon value propagation, we leverage the learned actor to train the value function backward from the terminal boundary using a windowed temporal curriculum, where each window is used as the boundary condition for the next window. Across benchmark problems up to 80 dimensions, our method learns accurate reachability value functions while improving stability over existing learning-based solvers. We further demonstrate observation-space scalability on F1-tenth racing with over 16,000-dimensional egocentric inputs. The learned safety filter generalizes zero-shot to unseen tracks and transfers to a physical RC car, achieving real-time robust collision avoidance.
Problem

Research questions and friction points this paper is trying to address.

Hamilton-Jacobi Reachability
High-Dimensional Nonlinear Systems
Curse of Dimensionality
Spatial Value Gradients
Reinforcement Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gradient-Free Neural Reachability
Bang-Bang Structure
Supervised Policy Learning
Windowed Temporal Curriculum
Safety-Critical Control
💼 Related Jobs
No related jobs found.
Zeyuan Feng
Zeyuan Feng
Stanford University
roboticsoptimal controlgame theorydeep learning
A
Ali Fuat Sahin
Department of Mechanical Engineering, École Polytechnique Fédérale de Lausanne, Switzerland
S
Santiago Thorup
Department of Aeronautics and Astronautics, Stanford University, United States
Somil Bansal
Somil Bansal
Assistant Professor, Stanford University
RoboticsArtificial intelligenceDynamic systems and control