TAKE 85: Testing Audiovisual filmmaKer's intEnt across 85 Hours of Film

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过创建TAKE 85基准,评估多模态大语言模型对导演意图的理解能力,揭示了现有模型在理解创作决策意图上的不足。
📝 Abstract
Films communicate through deliberate creative choices, including lighting, color, composition, editing, dialogue, music, and sound. Humans naturally interpret these signals as directorial intent, yet current multimodal large language models (MLLMs) are evaluated almost exclusively on understanding what happens rather than why it is presented that way. We introduce TAKE 85, the first benchmark for directorial-intent understanding, comprising 398 short films (85 hours) with expert-verified question-answer pairs spanning global and fine-grained visual and audio intent. Through controlled modality ablations, TAKE 85 enables systematic evaluation of multimodal reasoning. Experiments on state-of-the-art MLLMs reveal a substantial gap between perceptual recognition and intentional understanding: while models accurately describe events and narratives, they consistently fail to infer the communicative role of filmmaking decisions. Our results establish directorial intent as a previously overlooked dimension of multimodal understanding: even the strongest model reaches only 58 out of 100, and our ablations show that no input modality is sufficient on its own. All code, Q&As, and models are publicly available from https://github.com/KaiShinozakiConefrey/Take-85
Problem

Research questions and friction points this paper is trying to address.

directorial intent
multimodal large language models
filmmaking decisions
multimodal reasoning
perceptual recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

TAKE 85
directorial intent
multimodal reasoning
modality ablations
🔎 Similar Papers
No similar papers found.
K
Kaishuu Shinozaki-Conefrey
LIX, Ecole Polytechnique, IP Paris, Palaiseau, France; New York University, New York, NY, USA
O
Olivier Pascaud
LIX, Ecole Polytechnique, IP Paris, Palaiseau, France; Ecole nationale supérieure Louis-Lumière, Saint-Denis, France
Robin Courant
Robin Courant
École Polytechnique, IP Paris
Computer visionCinematographyDeep learning
X
Xi Wang
LIX, Ecole Polytechnique, IP Paris, Palaiseau, France
Dimitris Samaras
Dimitris Samaras
Stony Brook University
Computer VisionMachine LearningComputer GraphicsMedical Imaging
Vicky Kalogeiton
Vicky Kalogeiton
École Polytechnique, IP Paris
Computer VisionDeep Learning