Membership Inference in Fine-tuned Diffusion Language Models via Token-level Memorization Asymmetry

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过分析扩散语言模型中的token级记忆不对称现象,提出Q-Skew指标以进行成员推理攻击,揭示了该类模型的隐私风险。
📝 Abstract
Diffusion language models (DLMs) have recently emerged as an alternative modeling paradigm to autoregressive LMs, offering advantages such as parallel generation and bidirectional context modeling. Despite growing interest in their generative capabilities, the privacy risks of DLMs remain underexplored. We identify a phenomenon termed token-level memorization asymmetry through theoretical analysis of diffusion training dynamics. Building on this finding, we propose Q-Skew, a quantile-weighted skewness-based indicator for membership inference on finetuned DLMs. Experiments across multiple fine-tuning datasets and models show that our method outperforms existing baselines. Moreover, we show that Q-Skew can also facilitate other privacy violations, such as PII extraction. Our findings reveal a previously underexplored privacy attack surface and highlight the need for systematic privacy evaluation of DLMs.
Problem

Research questions and friction points this paper is trying to address.

Membership Inference
Diffusion Language Models
Token-level Memorization Asymmetry
Privacy Risks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Token-level Memorization Asymmetry
Q-Skew
Membership Inference