Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对个性化TTS模型的隐私泄露问题,提出了首个适用于TTS模型的黑盒成员推断攻击框架,通过优化查询生成和表示工程方法来识别成员身份。
📝 Abstract
Text-to-Speech (TTS) foundation models are increasingly fine-tuned on private datasets to synthesize highly personalized voices, introducing severe privacy risks by exposing both biometric identities and sensitive speech content. Existing black-box membership inference attacks (MIAs) follow a two-stage pipeline of query generation and representation engineering, both of which face unique challenges when adapted to TTS. For query generation, dual conditioning on synthesis text and reference speech creates a large and underexplored query design space with no established criterion for identifying an effective query. For representation engineering, the multi-level speech characteristics and temporal variability of speech make low-level representations and direct comparisons inadequate for capturing membership signals. To address these challenges, we present the first black-box MIA framework explicitly tailored to TTS models at both the speaker and record levels. For query generation, we characterize the feasible query space and establish two criteria, scorable extent and memorization elicitation, for evaluating five representative queries, identifying recitation as the strongest. For representation engineering, we obtain multi-level speech representations from embedding models and temporally align the generated and target audio for fine-grained comparison. Evaluations across three state-of-the-art TTS models (CosyVoice2, F5-TTS, and XTTS-v2) fine-tuned on two benchmark datasets (VCTK and British Dialect) reveal severe privacy leakage: speaker-level AUC remains above 0.80 and approaches 1.0 in the strongest settings, while record-level AUC ranges from 0.80 to 0.90 and remains effective even in challenging scenarios where both members and non-members are of the same speakers. We further identify speech characteristics associated with disproportionate vulnerability to memorization.
Problem

Research questions and friction points this paper is trying to address.

Black-box Membership Inference Attacks
Text-to-Speech Models
Query Generation
Representation Engineering
Privacy Risks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Black-Box Membership Inference Attacks
Text-to-Speech Models
Query Generation
Representation Engineering
Privacy Leakage
🔎 Similar Papers
No similar papers found.