Gradient Mirage: Trainable yet Label-Unidentifiable Gradients in Large Language Model Split Learning

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大型语言模型分裂学习中的梯度匹配攻击问题,提出Gradient Mirage方法,通过引入不一致性保护隐私同时保持优化效用。
📝 Abstract
Gradient matching attacks (GMAs) in LLM split learning (SL) rely on a critical yet underexplored assumption: the gradient exposed at the split interface is a faithful derivative of the client's full-label training objective. This gradient-objective consistency allows a curious server to recover private labels by searching for a sequence whose induced gradient explains the observation. We propose Gradient Mirage, a defense that breaks this consistency without discarding the optimization utility of the backward signal. Our key idea is to induce the adversary to solve a misspecified inverse problem, in which no plausible label sequence in the sequence space can explain the observed gradients. Concretely, Gradient Mirage achieves this by inducing inconsistency across three dimensions: objective, direction, and scale. Selective Autoregressive Supervision derives the exposed gradient from a masked surrogate loss rather than the full-label objective assumed by the attacker; Scale Blinding then applies randomized multiplicative rescaling, obscuring the gradient's natural magnitude; and Directional Privatization further randomizes the gradient direction while preserving its magnitude through the von Mises-Fisher (vMF) mechanism under a directional metric differential privacy guarantee. Crucially, utility is preserved: the Top segment still learns from all target tokens via Dual-Track Backpropagation, the exposed gradient remains informative since each supervised token retains its complete autoregressive context, and Bottom-Gradient Recovery restores the effective gradient for Bottom-segment optimization. Extensive experiments show that Gradient Mirage provides substantially stronger protection than existing defenses under comparable fine-tuning performance, achieving a better privacy-utility trade-off.
Problem

Research questions and friction points this paper is trying to address.

Gradient Matching Attacks
Large Language Model Split Learning
Privacy Leakage
Gradient-Objective Consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gradient Mirage
Selective Autoregressive Supervision
Scale Blinding
Directional Privatization
Dual-Track Backpropagation
🔎 Similar Papers
S
Shiyu Miao
State Key Laboratory for Novel Software Technology, Nanjing University
Yunlong Mao
Yunlong Mao
Nanjing University
SecurityPrivacyMachine Learning
Z
Zirui Huang
State Key Laboratory for Novel Software Technology, Nanjing University
L
Liang Yao
College of Computer Science and Software Engineering, Hohai University
T
Tianshuo Zheng
School of Mathematics, Nanjing University
Y
Yanhui Gu
Nanjing Normal University
Fan Liu
Fan Liu
Hohai University
computer vision
Sheng Zhong
Sheng Zhong
Nanjing University
computer networkssecurity and privacytheory of computing