PersonaShot: Benchmarking Person-Centric Narrative Continuity in Multi-Shot Video Generation

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of evaluation benchmarks for human narrative coherence in multi-shot video generation by constructing the first human-centric dataset comprising over one thousand clips and sixteen metrics. Through multimodal knowledge distillation and hierarchical temporal modeling, we propose a lightweight specialized evaluator aligned with human experts to precisely quantify cross-shot consistency in physical and emotional states. Our analysis reveals significant deficiencies in state-of-the-art models regarding narrative continuity. Demonstrating high correlation with expert judgment, this evaluator provides a systematic profiling of current limitations and establishes clear directions for enhancing the narrative capabilities of video generation systems.
📝 Abstract
Video generation is rapidly evolving from single-shot clips to multi-shot narratives, where the human character serves as the core narrative anchor. However, existing benchmarks mainly assess character appearance or individual-shot quality, without measuring whether physical and emotional states remain coherent across cuts. They also rarely provide criterion-specific evaluation methods, although physical continuity, facial dynamics, and cinematic relations require different visual, temporal, and relational evidence. To address these limitations, we introduce PersonaShot, the first person-centric benchmark for narrative continuity in multi-shot video generation. PersonaShot contains approximately 1,000 multi-shot segments and 16 metrics spanning physical continuity, affective dynamics, and cinematic grammar. \textbf{\textit{1)} Narrative Continuity Benchmark:} We evaluate character coherence across three temporal levels: within-shot states, cross-shot transitions, and sequence-level trajectories. \textbf{\textit{2)} Human-Aligned Specialist Evaluators:} We distill reasoning from a large multimodal teacher into lightweight criterion-specific evaluators, each grounded in the visual, temporal, or relational evidence required by its metric, and align them with expert human judgments. \textbf{\textit{3)} Systematic Evaluation and Insights:} Our evaluation reveals distinct capability profiles across state-of-the-art models and a clear gap between perceptual quality and cross-shot narrative continuity. Even visually compelling videos frequently exhibit physical-state resets, abrupt affective shifts, and broken cinematic relations across shots. Human studies further demonstrate strong agreement between our evaluators and expert judgments.
Problem

Research questions and friction points this paper is trying to address.

Multi-Shot Video Generation
Narrative Continuity
Person-Centric Evaluation
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Narrative Continuity Benchmark
Person-Centric Evaluation
Human-Aligned Specialist Evaluators
Multi-Shot Video Generation
Knowledge Distillation
🔎 Similar Papers
2024-05-22Annual Meeting of the Association for Computational LinguisticsCitations: 2