InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model

📅 2026-03-08
🏛️ arXiv.org
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出InfoMamba模型,通过结合线性过滤层与选择性循环流来解决序列建模中局部细节和长距离依赖捕捉之间的平衡问题。
📝 Abstract
Balancing fine-grained local modeling with long-range dependency capture under computational constraints remains a central challenge in sequence modeling. While Transformers provide strong token mixing, they suffer from quadratic complexity, whereas Mamba-style selective state-space models (SSMs) scale linearly but often struggle to capture high-rank and synchronous global interactions. We present a consistency boundary analysis that characterizes when diagonal short-memory SSMs can approximate causal attention and identifies structural gaps that remain. Motivated by this analysis, we propose InfoMamba, an attention-free hybrid architecture. InfoMamba replaces token-level self-attention with a concept bottleneck linear filtering layer that serves as a minimal-bandwidth global interface and integrates it with a selective recurrent stream through information-maximizing fusion (IMF). IMF dynamically injects global context into the SSM dynamics and encourages complementary information usage through a mutual-information-inspired objective. Extensive experiments on classification, dense prediction, and non-vision tasks show that InfoMamba consistently outperforms strong Transformer and SSM baselines, achieving competitive accuracy-efficiency trade-offs while maintaining near-linear scaling.
Problem

Research questions and friction points this paper is trying to address.

sequence modeling
local modeling
long-range dependency
computational constraints
attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

attention-free
concept bottleneck
information-maximizing fusion (IMF)
selective state-space models (SSMs)
💼 Related Jobs
No related jobs found.
Y
Youjin Wang
Renmin University of China, Beijing, China
J
Jiaqiao Zhao
University of Macau, Macau SAR, China
R
Rong Fu
Central South University, Changsha, China
R
Run Zhou
Renmin University of China, Beijing, China
Ruizhe Zhang
Ruizhe Zhang
Purdue University
Quantum computingOptimizationMachine learningComplexity theory
J
Jiani Liang
Renmin University of China, Beijing, China
S
Suisuai Cao
Central South University, Changsha, China
Feng Zhou
Feng Zhou
Associate Professor, Renmin University of China
statistical machine learningBayesian methodstochastic processAI for science