Designing Compact Neural Architectures via Neuron Gating and Mixed Activation

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational expense and trade-offs between compactness and performance in neural architecture search by proposing a general bilevel optimization framework. Through three scalable continuous relaxation formulations, discrete decisions are transformed into differentiable problems, enabling efficient architecture search at both neuron and activation levels. The proposed method outperforms DARTS on MNIST and CIFAR-10 benchmarks, achieving high accuracy with significantly fewer parameters. These results effectively validate its advantages in model compression and performance enhancement, establishing a new paradigm for constructing lightweight yet high-performance neural networks.
📝 Abstract
Neural Architecture Search (NAS) is naturally formulated as a bilevel optimization problem, where the upper-level optimizes the architecture using validation performance and the lower-level trains network parameters using training loss. However, NAS is computationally expensive due to discrete architectural decisions, exponentially growing search spaces, and the high cost of training candidate architectures. This work develops a general bilevel optimization framework for NAS across diverse architectures, including MLPs, CNNs, RNNs, and Transformers, to identify compact architectures with strong predictive performance. We propose three scalable formulations that replace discrete neuron- and activation-level decisions with continuous relaxations, enabling differentiable optimization over otherwise combinatorial architecture spaces. These formulations give rise to three NAS methods: NAS based on Neuron Gating (NAS-NG), NAS based on Mixed Activation (NAS-MA), and NAS based on Neuron Gating and Mixed Activation (NAS-NGMA). Experiments on MLPs and CNNs using MNIST and CIFAR-10 show that the proposed methods consistently identify compact architectures with competitive or improved predictive performance. On MNIST, NAS-NGMA achieves 98.68% test accuracy with 7.69M MLP parameters, while NAS-NG achieves 99.63% accuracy with only 0.26M CNN parameters. On CIFAR-10, the proposed methods consistently outperform vanilla DARTS. Further experiments demonstrate that NAS-NG can optimize substantially over-parameterized and literature-optimal architectures, improving accuracy while reducing parameters. These results establish relaxed bilevel optimization as a scalable alternative to discrete NAS and provide a general framework for efficient neuron- and activation-level architecture optimization.
Problem

Research questions and friction points this paper is trying to address.

Neural Architecture Search
Bilevel Optimization
Compact Neural Architectures
Computational Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Differentiable NAS
Bilevel Optimization
Neuron Gating
Mixed Activation
Continuous Relaxation
🔎 Similar Papers
No similar papers found.
A
Abhishek Shukla
Department of Management Sciences, IIT Kanpur, India
A
Ankur Sinha
Krishnamurthy Tandon School of AI, IIM Ahmedabad, India
F
Faiz Hamid
Department of Management Sciences, IIT Kanpur, India