Event-Driven Language Models with Sparse Neural Activity for Neuromorphic Hardware

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种方法,通过在重度量化线性注意力模型中诱导稀疏神经活动,减少计算量,解决大型语言模型推理时的内存限制和高计算成本问题。
📝 Abstract
Inference with transformer-based large language models (LLMs) is often limited by the memory-bound KV cache and quadratic attention cost. State-space models (SSMs) mitigate this through linear attention and fixed-size recurrent states, but their large dense linear projections remain computationally expensive even after quantization. We introduce a method that induces sparse neural activity in heavily quantized linear-attention models with minimal performance loss. Activations below a per-projection trainable threshold ($\pm Δ$) are nullified while preserving crucial outliers, achieving comparable performance to dense models with up to 4$\times$ fewer effective arithmetic operations. Targeting a multi-core, multi-chip neuromorphic platform, where event-driven execution converts unstructured sparsity into throughput at both the compute and communication levels, a capability GPU architectures fundamentally lack, we project up to 37$\times$ higher throughput and 16$\times$ lower power versus edge GPU inference of a comparable transformer-based model, and up to 5.4$\times$ improvements over the non-sparsified baseline. These results position sparse, quantized linear-attention models as a natural fit for deploying LLMs on event-driven multi-core platforms.
Problem

Research questions and friction points this paper is trying to address.

Event-Driven
Sparse Neural Activity
Neuromorphic Hardware
Large Language Models
Linear Attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

sparse neural activity
quantized linear-attention models
event-driven execution
neuromorphic hardware
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Simon Richter
Simon Richter
Department of Electrical and Computer Engineering, Aarhus University, Aarhus, Denmark
R
Ruhai Lin
Department of Computer Science and Engineering, University of California, Santa Cruz, Santa Cruz, USA
Jason Yik
Jason Yik
Harvard University
Taylor Kergan
Taylor Kergan
University of California, Santa Cruz
Machine Learning
Rui-Jie Zhu
Rui-Jie Zhu
Ph.D. Student, University of California, Santa Cruz
Brain-Inspired EngineeringLanguage Modeling
F
Farshad Moradi
Department of Electrical and Computer Engineering, University of Southern Denmark, Odense, Denmark
J
Jason Eshraghian
Department of Computer Science and Engineering, University of California, Santa Cruz, Santa Cruz, USA