ACE: Adapter Consolidation across Experts for Parameter-Efficient Fine-Tuning of MoE LLMs

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究针对MoE模型参数高效微调中适配器碎片化问题,提出ACE方法,通过整合冗余专家的适配器并执行分组计算,提高准确率和训练速度。
📝 Abstract
Parameter-efficient fine-tuning (PEFT) of mixture-of-experts (MoE) models commonly attaches a separate low-rank adapter to each expert. This expert-wise design fragments adaptation in three ways: capacity is split across narrow low-rank updates, gradient supervision becomes sparse and imbalanced under sparse routing, and execution is decomposed into many small GEMMs. We find that such expert-wise separation is often unnecessary, as subsets of LoRA adapters become functionally similar during fine-tuning, revealing redundancy among expert-specific adapters. Based on this redundancy, we propose ACE (Adapter Consolidation across Experts), which groups redundant experts and replaces their expert-specific adapters with group-shared higher-rank LoRA modules under the same PEFT budget. ACE further introduces grouped adapter execution, which consolidates fragmented expert-wise adapter computations into fewer, larger group-level GEMMs. Across evaluations covering 12 datasets and four MoE backbones, ACE achieves the highest observed mean accuracy among the parameter-matched PEFT methods on the three backbones with complete baseline coverage, while providing $1.31\times$ to $1.48\times$ wall-clock training speedup over expert-wise LoRA without increasing peak memory. Our code is available at https://github.com/UbiquitousAILab/ACE.
Problem

Research questions and friction points this paper is trying to address.

Parameter-efficient Fine-tuning
Mixture-of-Experts
Adapter
Redundancy
Fragmentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adapter Consolidation
Parameter-Efficient Fine-Tuning
Mixture-of-Experts
Grouped Adapter Execution
Higher-Rank LoRA