Mixture of Experts Softens the Curse of Dimensionality in Operator Learning
To address the computational and memory bottlenecks imposed by the curse of dimensionality in high-dimensional operator learning, this paper proposes the Mixture-of-Experts Neural Operator (MoNO): a framework that decomposes a global nonlinear operator into multiple lightweight expert sub-operators, routed input-adaptively via a learnable decision-tree mechanism. Theoretically, we establish the first distributed universal approximation theorem, proving that MoNO uniformly approximates any Lipschitz-continuous nonlinear operator in Sobolev spaces; each expert’s depth, width, and rank scale as O(ε⁻¹), ensuring controllable memory footprint compatible with standard hardware. We further derive the first quantitative approximation rate for classical neural operators. Experiments and theory jointly demonstrate that MoNO achieves ε-accuracy with significantly reduced complexity, overcoming both expressive and deployability limitations inherent to monolithic neural operators.