NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

πŸ“… 2026-08-13
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing video understanding benchmarks struggle to evaluate models’ ability to jointly comprehend narrative progression and culturally embedded meanings in high-context, non-English long-form videos. To address this gap, this work introduces the first large-scale benchmark for Japanese ultra-long videos, comprising 155 videos (146.8 hours) and 1,481 questions, uniquely integrating narrative tracking with cultural reasoning. The authors propose a hierarchical memory-driven annotation pipeline, a task-oriented question synthesis mechanism, and a two-stage native-speaker validation protocol coupled with iterative shortcut elimination to ensure cognitive depth and question quality. Evaluations reveal that current multimodal large language models exhibit significant limitations in long-range narrative integration and cultural inference.
πŸ“ Abstract
Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high-context, non-English media. To address this gap, we introduce NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video. NARU consists of 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. To construct the benchmark at this scale, we propose a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal. The construction process includes two native-speaker verification stages involving 68 annotators. Evaluations across eight model configurations reveal substantial limitations in both long-range narrative integration and culturally grounded reasoning. By exposing these persistent gaps, NARU offers a systematic testing ground for developing MLLMs capable of reliably interpreting long-form, high-context video.
Problem

Research questions and friction points this paper is trying to address.

long-form video understanding
narrative evolution
cultural nuance
high-context media
Japanese video
Innovation

Methods, ideas, or system contributions that make the work stand out.

narrative evolution
cultural nuance understanding
long-form video benchmark
hierarchical memory-based annotation
multimodal large language models