The Hamilton-Jacobi Theory of Deep Learning

📅 2026-05-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work establishes a precise mathematical correspondence between deep neural network training and inference and the theory of partial differential equations, interpreting training as solving an initial-value problem for the Hamilton–Jacobi equation. Each gradient update step is equivalent to selecting an initial condition for a viscous Hamilton–Jacobi equation and optimally fitting data via the Hopf–Cole propagator, while inference corresponds to evaluating the solution at specific points. By introducing a single deformation parameter ε, the framework unifies four perspectives—Hamilton–Jacobi PDEs, tropical geometry, convex optimization, and network architecture—for the first time. It yields a minimax optimal generalization rate of O(n⁻¹⁄⁽ᵈ⁺²⁾) at fixed time, reveals how ε governs adversarial robustness, provides an O(N) closed-form influence function, and characterizes the fold bifurcation of the entropy landscape induced by varying ε.
📝 Abstract
In this paper, training a neural network is identified, exactly, as a search through Hamilton--Jacobi initial-value problems: each gradient step selects the initial data of a viscous Hamilton--Jacobi equation whose Hopf--Cole propagator best fits the observations; at inference, the input is the spatial point at which that solution is evaluated and the initial condition is already encoded in the weights. The correspondence is exact for log-sum-exp layers and structural for broader architectures: residual networks, transformers, and recurrent architectures (RNNs, LSTMs, SSMs) each discretize the same class of Hamilton--Jacobi equations, with architecture-dependent Hamiltonian and viscosity. A single deformation parameter $\varepsilon$ unifies all four perspectives (network, tropical algebra, viscous PDE, convex optimization) in a commutative diagram closed under Lipschitz conditions. Quantitative consequences include: the minimax optimal generalization rate $O(n^{-1/(d+2)})$ for fixed $t$; adversarial robustness controlled by $\varepsilon$; backpropagation as the co-state equation of the Hamiltonian system for residual networks (Pontryagin Maximum Principle); scaling exponents consistent with data intrinsic dimension via PDE quadrature; and a closed-form $O(N)$ influence function (softmax attribution weights $π_j$) whose entropy landscape undergoes fold bifurcations as $\varepsilon$ increases, each merging attribution basins.
Problem

Research questions and friction points this paper is trying to address.

Hamilton-Jacobi equation
deep learning
neural network training
viscous PDE
initial-value problem
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hamilton-Jacobi equation
viscous PDE
Hopf-Cole transform
neural architecture unification
influence function
🔎 Similar Papers
2024-07-25arXiv.orgCitations: 0
Jose Marie Antonio Miñoza
Jose Marie Antonio Miñoza
Senior Data Scientist, Center for AI Research; System Modeling and Simulation Lab, UP Diliman;
mathematical modelingoptimizationmachine learningscientific ml
E
Erika Fille T. Legara
Center for AI Research PH, Asian Institute of Management
C
Christopher P. Monterola
Asian Institute of Management