VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the saturation of existing video understanding benchmarks and their inadequacy in evaluating agent capabilities by constructing a novel benchmark tailored for general-purpose AI assistants. We introduce the first multi-turn, tool-augmented evaluation framework that transcends single-turn QA limitations, encompassing 271 complex real-world tasks with quality assured through human-AI collaboration and triple expert verification. This benchmark facilitates a paradigm shift from traditional video understanding toward an agentic approach. Experimental results demonstrate that even state-of-the-art models achieve accuracy below 60%, effectively revealing significant deficiencies in current models regarding complex video reasoning and tool interaction. Collectively, this work establishes a rigorous standard for assessing next-generation video agents and highlights critical challenges requiring further investigation in embodied AI research.
📝 Abstract
Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME leaderboard, suggesting that conventional single-turn video understanding tasks are becoming increasingly saturated and insufficient for assessing the intelligence of advanced MLLMs. Towards this end, we introduce VideoGAIA, an agentic video understanding benchmark for general artificial intelligence (AI) assistants. Moving beyond one-shot video question answering, VideoGAIA formulates video understanding as a multi-turn, tool-augmented interaction process, where models must iteratively perceive videos, invoke external tools, gather complementary information, and integrate multimodal evidence across turns. VideoGAIA contains 271 model-human co-designed tasks covering diverse and complex real-world scenarios. Each video-question-answer instance is independently verified by three human experts to ensure both correctness and appropriate difficulty. All evaluated MLLMs, including frontier models such as GPT-5.5 and Kimi-K3, achieve less than 60% accuracy on VideoGAIA, highlighting its value as a high-quality and timely benchmark for evaluating next-generation MLLMs. We hope that VideoGAIA will facilitate the transition from conventional video understanding toward agentic video understanding.
Problem

Research questions and friction points this paper is trying to address.

Agentic Video Understanding
Multimodal Large Language Models
Benchmark Saturation
Multi-turn Interaction
Tool-augmented Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Video Understanding
Tool-Augmented Interaction
Multi-turn Reasoning
Benchmark
Multimodal Large Language Models
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30