PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决文本到音视频生成中音频评估不足的问题,提出PRISM-Bench基准,通过4个感知维度和35个细粒度标准对音频进行评测。
📝 Abstract
Text-to-audio-video (T2AV) generation has advanced rapidly, but its evaluation still underestimates the audio modality. Existing benchmarks either treat audio as an auxiliary component of video quality or assess it in isolation from audiovisual grounding, making it difficult to diagnose where current systems truly succeed or fail in audio generation. We present PRISM-Bench, the first audio-centric diagnostic benchmark for T2AV generation. Built from a rigorously curated dataset of 900 human-verified samples, PRISM-Bench factorizes audio evaluation along two orthogonal axes: audio type (Speech, Music, and Sound) and sound-source visibility (On-screen vs. Off-screen). It evaluates generated content across four perceptual dimensions (Audio-Visual Coherence, Audio Quality, Audio Expressiveness, and Prompt Following) with 35 fine-grained criteria. To ensure reliable assessment, we adopt an enhanced MLLM-as-a-Judge protocol based on blind, side-by-side comparison against ground-truth references, demonstrating strong alignment (over 70% mean agreement) with human raters. Our evaluation of recent T2AV systems highlights a significant performance gap between frontier and open-source models. Furthermore, we demonstrate that current generation paradigms overfit to perceptual fidelity while struggling with complex grounding and control tasks, particularly in generating music and synchronized On-screen audio.
Problem

Research questions and friction points this paper is trying to address.

Text-to-audio-video
audio evaluation
audio-visual coherence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audio-Centric Benchmark
Text-to-Audio-Video Generation
Perceptual Dimensions
Enhanced MLLM-as-a-Judge Protocol
Human-Verified Samples
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30
Y
Yuchen Sun
Shanghai Artificial Intelligence Laboratory
Qian Yang
Qian Yang
Meituan, Intern@Alibaba Qwen, Zhejiang University
SpeechAudio
J
Jun Wang
Meituan
Detai Xin
Detai Xin
The University of Tokyo
Speech processingSpeech synthesisMachine Learning
G
Guoqiao Yu
Meituan
G
Guanglu Wan
Meituan
Q
Qi Jia
Shanghai Artificial Intelligence Laboratory