NeuroFlex: Lossless Element-Level ANN-SNN Co-Execution for Efficient Sparse Inference

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
NeuroFlex通过在无精度损失情况下独立为每个输出元素分配ANN或SNN执行模式,解决了稀疏推理中能量和延迟效率低下的问题。
📝 Abstract
Sparse DNN accelerators specialize in ANN or SNN execution, leaving energy or latency on the table when workload characteristics vary within a layer. Hybrid accelerator designs that switch modes at layer or tile granularity suffer from low PE utilization since one core type idles whenever the other is active. NeuroFlex is the first accelerator to assign every output element independently to ANN or SNN execution mode with zero accuracy loss. We extend integer-exact ANN-SNN equivalence from layers to individual output elements, thereby enabling mode switching with no conversion error. An offline cost-guided scheduler scores each element by its marginal energy-delay trade-off and packs work across PEs, achieving 97-99% PE utilization compared to 40-45% for layer-wise hybrids. NeuroFlex reduces EDP by 57-67% over a strong ANN-only baseline and delivers up to 2.5x speedup over a dual-sparse SNN-only baseline. Our cost-guided scheduler improves throughput by 16-19% over random element assignment across vision, language, and transformer workloads.
Problem

Research questions and friction points this paper is trying to address.

Sparse DNN accelerators
ANN or SNN execution
PE utilization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lossless Element-Level Co-Execution
Cost-Guided Scheduler
Energy-Delay Trade-off
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
V
Varun Manjunath
Electrical Engineering, Indian Institute of Technology Madras, Chennai, India
P
Pranav Ramesh
Computer Science and Engineering, Indian Institute of Technology Madras, Chennai, India
Gopalakrishnan Srinivasan
Gopalakrishnan Srinivasan
Assistant Professor at IIT Madras
RISC-V SoCAI Accelerator ArchitecturesDeep LearningSpiking Neural Networks