On the Approximation Power of SiLU Networks: Exponential Rates and Depth Efficiency

📅 2025-12-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work investigates the approximation power of SiLU-activated neural networks for smooth functions, aiming to overcome the polynomial convergence rate limitation inherent in ReLU networks. We propose a hierarchical construction based on efficient quadratic approximations, achieving— for the first time with constant depth (O(1))—an approximation error decay of O(ω⁻²ᵏ). We establish tight depth–size trade-offs for approximating Sobolev functions, attaining parameter complexity O(ε⁻ᵈ⁄ⁿ), which significantly improves upon comparable ReLU-based constructions. Theoretically, we prove that SiLU networks achieve exponential approximation rates while maintaining constant depth and optimal size scaling. This is the first systematic theoretical characterization revealing SiLU’s intrinsic advantage in approximating smooth functions, providing rigorous justification for activation function selection in deep learning.

Technology Category

Application Category

📝 Abstract
This article establishes a comprehensive theoretical framework demonstrating that SiLU (Sigmoid Linear Unit) activation networks achieve exponential approximation rates for smooth functions with explicit and improved complexity control compared to classical ReLU-based constructions. We develop a novel hierarchical construction beginning with an efficient approximation of the square function $x^2$ more compact in depth and size than comparable ReLU realizations, such as those given by Yarotsky. This construction yields an approximation error decaying as $mathcal{O}(ω^{-2k})$ using networks of depth $mathcal{O}(1)$. We then extend this approach through functional composition to establish sharp approximation bounds for deep SiLU networks in approximating Sobolev-class functions, with total depth $mathcal{O}(1)$ and size $mathcal{O}(varepsilon^{-d/n})$.
Problem

Research questions and friction points this paper is trying to address.

Exponential approximation rates for smooth functions
Efficient depth and size compared to ReLU networks
Sharp bounds for Sobolev-class function approximation
Innovation

Methods, ideas, or system contributions that make the work stand out.

SiLU networks achieve exponential approximation rates for smooth functions
Novel hierarchical construction efficiently approximates square function with compact depth
Deep SiLU networks approximate Sobolev functions with shallow depth and controlled size
K
Koffi O. Ayena
Université de Lomé, Laboratoire de Modélisations Mathématiques et Applications, Lomé, Togo; Université de Belfort de Montbéliard, Laboratoire Interdisciplinaire Carnot de Bourgogne, ICB/UTBM, UMR 6303 CNRS, Belfort, France