π€ AI Summary
Current evaluations of instruction-following in text-to-music models are susceptible to output prior bias, making it difficult to discern whether generated attributes stem from the given instructions or the modelβs inherent preferences. This work proposes a matched counterfactual evaluation framework that disentangles instruction controllability from output priors by comparing outputs from neutral prompts against those from swapped target prompts, under controlled conditions including shared random seeds, frozen adapters, and external discriminator validation. Integrating blind expert annotations with a multi-seed sentinel mechanism, this approach provides the first systematic quantification of model controllability over tonality and metrical grouping. Experiments reveal that ACE-Step 1.5 and Stable Audio 3 Medium exhibit significant tonal control, whereas LeVo2 does not; furthermore, the high consistency in quadruple-meter structures is primarily attributable to model priors rather than instruction adherence.
π Abstract
Prompted attribute agreement is widely used as evidence of text-to-music controllability, yet a requested attribute may occur simply because it is already common in the model's output distribution. We introduce a matched counterfactual evaluation that separates target occurrence from instruction-attributable control. Each family contains a neutral input that omits the scored attribute and two otherwise matched inputs that swap the requested target. All three are rendered through frozen native-interface adapters with a shared seed. Applied to global key and beat grouping in three open systems, this design changes the empirical conclusion. ACE-Step 1.5 and Stable Audio 3 Medium exhibit substantial key control, whereas LeVo2 does not. For beat grouping, the same models redirect toward the rare three-beat target, but high four-beat agreement is largely inherited from neutral outputs: Stable Audio 3 produces four-beat grouping in 0.97 of neutral cases but only 0.56 under its explicit four-beat treatment. Off-attribute placebos, external recognizer validation, blind expert annotation, and multi-seed sentinels support the attribution. When targets have unequal output priors, agreement describes what a model produced, while matched neutral and target-swap contrasts test whether the instruction changed it.