How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech
This study addresses the limited understanding of how natural language instructions modulate acoustic outputs in current stylized text-to-speech (TTS) systems, which hinders model controllability and failure attribution. For the first time, the diffusion attention attribution method (DAAM) is introduced to the speech generation domain to perform cross-attention attribution analysis across 25 network layers and 24 ODE steps of the CapSpeech-TTS model, enabling fine-grained visualization and quantification of the influence of style-descriptive words. The findings reveal that style words exert a global regulatory effect, with their attention intensity significantly correlated with fundamental frequency and energy. This influence is most pronounced in deeper network layers—particularly layer 17, which exhibits the strongest selectivity—and during early ODE integration steps.