CoRe-MoE: Compact Reusable MoE for Continual Multimodal Instruction Tuning

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决持续多模态指令调优中参数开销大的问题,提出CoRe-MoE框架,通过重用初始专家库中的方向基和训练紧凑坐标专家来提高效率。
📝 Abstract
Continual multimodal instruction tuning requires multimodal large language models to acquire new task abilities sequentially while preserving previously learned knowledge. LoRA-MoE provides a promising solution by introducing expert-based capacity, but repeatedly learning and maintaining full LoRA experts leads to substantial parameter overhead. This raises a natural question: is full expert expansion necessary for every new task? To answer it, we analyze the SVD of task-specific LoRA updates and observe substantial overlap in their input- and output-side LoRA direction subspaces, with task-specific adaptation largely captured by lightweight coordinates over these subspaces. Motivated by this observation, we propose CoRe-MoE, a Compact Reusable MoE framework for parameter-efficient continual multimodal instruction tuning. CoRe-MoE extracts reusable input- and output-side direction bases from an initial expert bank, and for subsequent tasks trains only compact coordinate experts together with task-specific low-rank routers. Experiments on two representative MLLMs show that CoRe-MoE improves final average performance over the strongest competing baseline by up to 5.90 points, while using less than 1% of the trainable parameters required by sequential LoRA for later tasks. The code is publicly available at https://github.com/runzezz/CoRe-MoE.
Problem

Research questions and friction points this paper is trying to address.

Continual Multimodal Instruction Tuning
LoRA-MoE
Parameter Overhead
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compact Reusable MoE
Continual Multimodal Instruction Tuning
Parameter-efficient
Task-specific Low-rank Routers
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Runze Liu
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
N
Naibin Gu
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
M
Mingxu Ai
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
Yuqing Li
Yuqing Li
East China Normal University
Deep Learning Theory
Peng Fu
Peng Fu
Institute of Information Engineering, Chinese Academy of Sciences
Natural Language Processing
Zheng Lin
Zheng Lin
Institute of Information Engineering, CAS
NLP
Weiping Wang
Weiping Wang
School of Information Science and Engineering, Central South University
Computer NetworkNetwork Security