The Hamilton-Jacobi Theory of Deep Learning
This work establishes a precise mathematical correspondence between deep neural network training and inference and the theory of partial differential equations, interpreting training as solving an initial-value problem for the Hamilton–Jacobi equation. Each gradient update step is equivalent to selecting an initial condition for a viscous Hamilton–Jacobi equation and optimally fitting data via the Hopf–Cole propagator, while inference corresponds to evaluating the solution at specific points. By introducing a single deformation parameter ε, the framework unifies four perspectives—Hamilton–Jacobi PDEs, tropical geometry, convex optimization, and network architecture—for the first time. It yields a minimax optimal generalization rate of O(n⁻¹⁄⁽ᵈ⁺²⁾) at fixed time, reveals how ε governs adversarial robustness, provides an O(N) closed-form influence function, and characterizes the fold bifurcation of the entropy landscape induced by varying ε.