🤖 AI Summary
This study addresses the absence of multi-output generalization bounds in multi-task deep learning by modeling network layers as Koopman operators and deriving Rademacher complexity bounds within a vector-valued RKHS framework. By decoupling output correlations from layer operator norms, we establish a cross-task shared operator learning framework alongside a finite-rank representation theorem. Furthermore, novel generalization bounds independent of smoothness indices are obtained under Brownian mechanisms. Theoretically, we derive specific bounds for invertible and width-expanded architectures. Empirical validation on synthetic datasets and MNIST confirms the effectiveness of the proposed complexity proxy. Collectively, this work provides a rigorous operator-theoretic foundation for analyzing generalization in multi-task learning scenarios, bridging the gap between deep network architecture and functional analysis.
📝 Abstract
We develop operator-theoretic generalization bounds for deep multi-output function classes by representing network layers as Koopman composition operators on vector-valued reproducing kernel Hilbert spaces. In vector-valued Sobolev RKHSs, we derive Rademacher complexity bounds for invertible and width-expanding injective architectures. The estimates separate the output-coupling contribution, represented by the trace of the task matrix, from the layerwise operator norms, Sobolev symbol ratios, determinant factors, and restriction constants generated by the linear maps. We then analyze a distinct one-dimensional Brownian/Cameron--Martin regime. Using the exact anchored derivative-norm characterization of the vector-valued Brownian RKHS, we obtain layerwise bounds for domain-preserving scalar linear maps and anchored diffeomorphic activations; the corresponding factors scale as $|W_l|^{1/2}$ and $\|σ_l'\|_\infty^{1/2}$, respectively, and do not involve Sobolev smoothness exponents. Because the Sobolev and Brownian results concern different hypothesis spaces, neither is asserted to dominate the other uniformly. We additionally formulate shared operator learning across tasks, prove a finite-rank representer theorem, derive the exact finite-dimensional problem for squared loss, and establish a target-transfer bound when the learned operator is obtained independently of the target sample. Synthetic and MNIST studies examine stabilized Sobolev-inspired and Brownian-inspired complexity proxies; these empirical proxies are not evaluations of the proved bounds for rank-deficient architectures.