RAPTOR: Role-Aware Private Training for Mixture-of-Experts

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
Differentially private (DP) fine-tuning methods treat sparse Mixture-of-Experts (MoE) models as a single dense block, ignoring that shared layers see all data while experts only see routed records. We identify and formally characterize three resulting failure modes: global clipping suppresses expert gradients, batch-level normalization dilutes sparse expert updates, and fixed privacy noise degrades signal-to-noise ratio on low-load experts. We introduce RAPTOR - a Role-Aware Private Training framework, which alternates shared and expert optimization and targets each failure directly, using expert-specific clipping and noise together with a public expected-owner denominator and a count-independent update schedule that avoids conditioning on private, realized expert counts. We prove the resulting mechanism satisfies $(\varepsilon,\delta)$-DP: because each record is assigned to exactly one owner expert, per-expert mechanisms within a layer compose in parallel, so updating all $E$ experts costs no more, in privacy terms, than updating one, with shared and expert streams composing sequentially across training. We further derive a bias-variance decomposition of the public-denominator estimator showing its bias grows predictably with routing imbalance, yielding a privacy-free rule for selecting which layer to protect from routing entropy measured on a small public corpus. Experiments on Switch Transformer and OLMoE fine-tuning across GLUE tasks, and on DeepSeek-VL2-Tiny, show consistent gains over standard DP baselines across several privacy levels ($\varepsilon$), with the largest margins typically at the tightest budgets. Code and models are publicly available: https://github.com/leduckhai/RAPTOR
Problem

Research questions and friction points this paper is trying to address.

Differentially Private
Mixture-of-Experts
Sparse Models
Gradient Clipping
Normalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Role-Aware Private Training
Mixture-of-Experts (MoE)
Differential Privacy (DP)
Expert-Specific Clipping
Public Expected-Owner Denominator
🔎 Similar Papers
No similar papers found.
D
Duc Dm
KAIST
Khai Le-Duc
Khai Le-Duc
University of Toronto
Artificial IntelligenceHeal the world
N
Nguyen Do
University of Florida
M
Minh Son Hoang
KAIST
Florent Draye
Florent Draye
PhD Student, MPI-IS
LLM Interpretability
T
Thai Hoang
Salesforce AI Research
H
Hoang Phuong Dam
KAIST
Jiarui Liu
Jiarui Liu
Carnegie Mellon University
Natural Language Processing
Chris Ngo
Chris Ngo
Knovel Engineering
Terry Jingchen Zhang
Terry Jingchen Zhang
ETH Zurich
(Multimodal) ReasoningAI SafetyActionable InterpretabilityAI4ScienceAstrophysics
A
Anh Le Duc Tran
Hanoi University of Science and Technology
N
Nhat Do Minh
Vietnam National University, Hanoi
M
Minh Ngoc Le
University of Toronto, Vector Institute
My T. Thai
My T. Thai
Professor, University of Florida, IEEE Fellow
Explainable AISecurity and PrivacyNetwork ScienceOptimization
Ran Xu
Ran Xu
Salesforce Research
computer visionmachine learningdata mining
Silvio Savarese
Silvio Savarese
Associate Professor of Computer Science at Stanford University
Computer vision
Mona Diab
Mona Diab
Professor & Director of Language Technologies Institute, Carnegie Mellon University, ACL Fellow
Responsible AINLP/CLArabic NLPCross lingual/multilingual & Low Resource Lang Processing
Bernhard Schölkopf
Bernhard Schölkopf
Director, Max Planck Institute for Intelligent Systems & ELLIS Institute Tübingen; Professor at ETH
Machine LearningCausal InferenceArtificial IntelligenceComputational PhotographyStatistics
Zhijing Jin
Zhijing Jin
Max Planck Institute
Natural Language ProcessingCausal InferenceMachine LearningArtificial IntelligenceLLMs
H
Huy L. Nguyen
Northeastern University
Daeyoung Kim
Daeyoung Kim
Professor of School of Computing, KAIST
Cloud ComputingInternet of ThingsMachine Learning