Finite Neural Networks as Mixtures of Gaussian Processes: From Provable Error Bounds to Prior Selection
Existing finite-width deep neural networks lack analytically tractable Gaussian process (GP) approximations with provable error bounds. Method: We propose the first Gaussian Process Mixture (GPM) approximation framework with certified error bounds: leveraging Wasserstein distance to model output distributions layer-wise, it achieves ε-accurate approximation of arbitrary non-i.i.d. parameterized networks over finite input sets. The method integrates optimal transport theory with hierarchical probabilistic modeling, yielding differentiable error bounds that guide network parameter optimization toward user-specified prior distributions. Results: Experiments demonstrate that GPM enables controllable-accuracy approximation on both regression and classification tasks, while simultaneously supporting principled uncertainty quantification and Bayesian prior design—bridging finite-width neural networks and rigorous GP inference with guaranteed approximation quality.