Let SSMs be ConvNets: State-space Modeling with Optimal Tensor Contractions

📅 2025-01-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
State space models (SSMs) underperform in audio tasks due to limited modeling capacity, high computational overhead, and lack of inductive bias from convolutional priors. Method: We propose Centaurus—a fully SSM-based architecture that reformulates generalized SSMs as learnable tensor contraction operations with automatically optimized contraction orders; it is the first to systematically integrate group convolution, full convolution, and bottleneck structures into a heterogeneous SSM design. Contribution/Results: Centaurus is the first high-performance automatic speech recognition (ASR) model built entirely upon SSMs—without LSTM, explicit CNNs, or attention mechanisms. It outperforms homogeneous SSM baselines of comparable size on keyword spotting, speech denoising, and ASR. On ASR, it achieves state-of-the-art (SOTA) performance while significantly reducing memory consumption and computational cost during both training and inference.

Technology Category

Application Category

📝 Abstract
We introduce Centaurus, a class of networks composed of generalized state-space model (SSM) blocks, where the SSM operations can be treated as tensor contractions during training. The optimal order of tensor contractions can then be systematically determined for every SSM block to maximize training efficiency. This allows more flexibility in designing SSM blocks beyond the depthwise-separable configuration commonly implemented. The new design choices will take inspiration from classical convolutional blocks including group convolutions, full convolutions, and bottleneck blocks. We architect the Centaurus network with a mixture of these blocks, to balance between network size and performance, as well as memory and computational efficiency during both training and inference. We show that this heterogeneous network design outperforms its homogeneous counterparts in raw audio processing tasks including keyword spotting, speech denoising, and automatic speech recognition (ASR). For ASR, Centaurus is the first network with competitive performance that can be made fully state-space based, without using any nonlinear recurrence (LSTMs), explicit convolutions (CNNs), or (surrogate) attention mechanism. Source code is available at github.com/Brainchip-Inc/Centaurus
Problem

Research questions and friction points this paper is trying to address.

State Space Model
Sound Information Processing
Resource Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

State Space Model
Neural Network Architecture
Speech Recognition
💼 Related Jobs
No related jobs found.
Y
Yan Ru Pei
Brainchip Inc., Laguna Hills, CA 92653, USA