Mamba-CL: Optimizing Selective State Space Model in Null Space for Continual Learning

๐Ÿ“… 2024-11-23
๐Ÿ›๏ธ arXiv.org
๐Ÿ“ˆ Citations: 1
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Catastrophic forgetting severely hinders long-term adaptability in continual learning. This paper proposes the first Mamba-based, forgetfulness-free fine-tuning framework for class-incremental continual learning. Our method maps historical task features into a subspace and applies orthogonal parameter updates within its nullspaceโ€”thereby preserving output consistency of the State Space Model (SSM) core across tasks. We theoretically derive and enforce consistency constraints on four time-invariant SSM parameters, simplifying both the recurrent structure and discretization procedure. Crucially, this work introduces nullspace projection to the Mamba architecture for the first time, enabling efficient, replay-free, and regularization-free continual learning. Evaluated on four standard class-incremental benchmarks, our approach consistently outperforms state-of-the-art methods. The implementation is publicly available.

Technology Category

Application Category

๐Ÿ“ Abstract
Continual Learning (CL) aims to equip AI models with the ability to learn a sequence of tasks over time, without forgetting previously learned knowledge. Recently, State Space Models (SSMs), particularly the Mamba model, have achieved notable success in computer vision. Building on the strengths of SSMs, this study explores leveraging the Mamba model for CL. Therefore, we introduce Mamba-CL, a framework that continuously fine-tunes the core SSMs of the large-scale Mamba foundation model by updating parameters orthogonal to the feature subspace of previous tasks. This approach theoretically guarantees the consistency objective aiming to preserves consistent output for each SSM module across both previous and current tasks, so as to overcome catastrophic forgetting issue. Specifically, we achieve this goal by deducing the overall consistency constraints on four key time-invariant parameters in the Mamba model, streamlining its recurrent state-space structure and non-linear discretization process in SSM. In practice, we apply the null-space projection to efficiently implement the orthogonality within Mamba model. Extensive experiments on four class-incremental benchmarks demonstrate the effectiveness of Mamba-CL for anti-forgetting, achieving superior performances to state-of-the-art methods. Code is available in the supplementary materials.
Problem

Research questions and friction points this paper is trying to address.

Optimizing Mamba model for continual learning tasks
Preventing catastrophic forgetting in sequential task learning
Ensuring consistent output across previous and current tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fine-tunes Mamba SSMs with orthogonal parameter updates
Ensures output consistency via time-invariant parameter constraints
Implements orthogonality using null-space projection
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
De Cheng
De Cheng
Associate Professor, Xidian University
Computer VisionDeep LearningMachine LearningData Compression
Y
Yue Lu
School of Computer Science, Northwestern Polytechnical University, China
L
Lingfeng He
School of Telecommunications Engineering, Xidian University, China
Shizhou Zhang
Shizhou Zhang
Northwestern Polytechnical University
computer visionmachine learning
X
Xi Yang
School of Telecommunications Engineering, Xidian University, China
Nannan Wang
Nannan Wang
Professor, Xidian University
Computer VisionMachine LearningPattern Recognition
X
Xinbo Gao
Chongqing University of Posts and Telecommunications, China