DVBench: Benchmarking MLLMs for Understanding Dynamic Charts and Narratives in Data Videos

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决动态图表和数据视频中结构化叙述理解的问题,研究者通过DVBench基准测试评估多模态语言模型,并分析了不同模型的表现及特点。
📝 Abstract
While MLLMs have made significant strides in chart comprehension and video understanding, current evaluations largely isolate these capabilities, leaving a critical gap in understanding temporally evolving structured visual information. To address this gap, we introduce DVBench, a benchmark for evaluating MLLMs on data videos, a storytelling medium that integrates dynamic charts with structured narratives. We decompose data video understanding into five dimensions. DVBench comprises 300 real-world data videos and 1,000 human-verified QA pairs curated through a rigorous semi-automated pipeline. Extensive evaluations of nine MLLMs show that Gemini-3.1-Pro achieves the best overall performance, while Kimi-k2.5 is the strongest open-source model. We further identify two notable phenomena: open-source model performance does not scale strictly with parameter size, and narrative proficiency does not guarantee visual capability. Fine-grained analyses and ablation studies further reveal dimension-specific weaknesses and the effects of frame configurations and subtitle inputs, informing future MLLM development. DVBench is publicly available at https://bomiaowang.github.io/DVBench/.
Problem

Research questions and friction points this paper is trying to address.

MLLMs
dynamic charts
data videos
temporally evolving structured visual information
narratives
Innovation

Methods, ideas, or system contributions that make the work stand out.

DVBench
data videos
dynamic charts
narrative proficiency
MLLMs evaluation
🔎 Similar Papers