Voice ''Cloning'' is Style Transfer

📅 2026-05-15

📈 Citations: 0

✨ Influential: 0

career value

185K/year

🤖 AI Summary

This study challenges the common misconception that mainstream voice cloning technologies faithfully replicate a speaker’s voice, demonstrating instead that they systematically introduce stylistic biases. Through comprehensive human subjective evaluations, acoustic feature analyses (e.g., accent and speech rate), and variance measurements in audio embedding spaces, the work reveals that cloned voices are consistently perceived as more authoritative, warmer, customer-service-like, and anthropomorphic. These perceptual shifts significantly increase user trust and willingness to disclose sensitive information. Concurrently, the diversity of vocal characteristics is markedly reduced, leading to voice homogenization that threatens speaker identity authenticity and compromises user behavioral security. This paper thus establishes that current voice cloning operates fundamentally as a form of style transfer rather than accurate vocal reproduction.

📝 Abstract

Artificially generated speech is increasingly embedded in everyday life. Voice cloning in particular enables applications where identity preservation is important, such as completing a recording, dubbing in a new language, or preserving the voices of individuals with speech loss. However, in our work, we find that despite the term, voice cloning does not faithfully ''clone'' an individual's voice. Instead, we find that widely-used voice cloning models systematically apply style transfer to source voices. As rated by human annotators, cloned voices are perceived as more authoritative, warm, customer-service-like, and human-like compared to their sources. Human annotators also report greater trust in cloned voices than source voices, and a greater willingness to disclose sensitive personal information to them. Our work furthermore shows that voice cloning leads to homogenization of speaker characteristics, as measured by reduced variance in accent, speaking rate, and the audio embedding space. Together, our results highlight a new set of limitations and risks of voice cloning technology and their potential impact on human behavior.

Problem

Research questions and friction points this paper is trying to address.

voice cloning

style transfer

speaker homogenization

human perception

speech synthesis

Innovation

Methods, ideas, or system contributions that make the work stand out.

voice cloning

style transfer

speaker homogenization