Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对多指令长视频编辑难题,提出结合大语言模型和视觉-语言模型的代理编辑框架,确保片段间一致性、多指令解耦及时空结构无损。
📝 Abstract
While generative AI has significantly advanced video editing, existing methods primarily focus on single-shot or short video clips. Editing long videos with multiple instructions remains a formidable challenge. Naive chunking strategies, e.g., fixed-duration segmentation, often lead to entity fragmentation, severe editing hallucinations, and disrupted temporal continuity. To bridge this gap, we introduce the Multi-Instruction Multi-Shot Long-Video Editing (MMLVE) task, which is structured around three core objectives: Cross-Shot Editing Consistency (CSEC), Multi-Instruction Decoupling (MID), and Zero-Destruction on Spatiotemporal Structure (ZDSS). To tackle these three unique challenges, we introduce an agentic editing framework that leverages the synergy of Large Language Models (LLMs) and Vision-Language Models (VLMs) to achieve shot-level video decoupling and precise instruction parsing. Furthermore, to comprehensively evaluate this task, we construct MMLVE-Bench, which is an MMLVE-focused dataset characterized by complex real-world spatiotemporal dynamics, high-density heterogeneous instructions, and sparse, random entity distributions. Three MMLVE-focused evaluation metrics are further exploited to assess the quality of the editing results. Extensive experiments demonstrate that our MMLVE-Agent outperforms existing closed-source SOTA approaches (e.g., Seedance 2.0), successfully eliminating editing hallucinations, preserving cross-shot editing consistency, and attaining seamless spatiotemporal transitions.
Problem

Research questions and friction points this paper is trying to address.

Multi-Shot Video Editing
Cross-Shot Editing Consistency
Multi-Instruction Decoupling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Instruction Multi-Shot Long-Video Editing
Agentic Editing Framework
Cross-Shot Editing Consistency
Large Language Models
Vision-Language Models
💼 Related Jobs
No related jobs found.
Chenyang Wu
Chenyang Wu
Ph.D. Candidate, LAMDA, Nanjing University
reinforcement learningartificial intelligence
Fuchen Long
Fuchen Long
University of Science and Technology of China
Video Analysis
B
Binyuan Huang
Smart Creation Platform Department, Online Video BU, Tencent
X
Xinlong Sun
Smart Creation Platform Department, Online Video BU, Tencent
X
Xi Chen
Smart Creation Platform Department, Online Video BU, Tencent
C
Chun-Le Guo
VCIP, CS, Nankai University
Chongyi Li
Chongyi Li
Professor, Nankai University
Computer VisionComputational ImagingComputational PhotographyUnderwater Imaging