MTAVG-Bench: A Comprehensive Benchmark for Evaluating Multi-Talker Dialogue-Centric Audio-Video Generation
Existing evaluation benchmarks struggle to effectively assess critical issues in multi-speaker conversational video generation, such as identity drift, unnatural turn-taking, and audio-visual asynchrony. This work proposes the first fine-grained audiovisual generation evaluation framework tailored to this scenario, introducing a comprehensive benchmark comprising 1.8K videos and 2.4K structured question-answer pairs, constructed via a semi-automatic pipeline. The framework evaluates models across four dimensions: audiovisual fidelity, temporal consistency, social interaction coherence, and cinematic expressiveness, enabling precise failure analysis and targeted model refinement. Experiments on twelve leading open- and closed-source models reveal that Gemini 3 Pro achieves the best overall performance, while certain open-source models excel in signal fidelity and temporal consistency.