Plans You Can Check: Verifier-Grounded Learning of an Open-Weight Planner for Executable Video-Editing

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过训练一个紧凑的开放权重规划器RefineCut来解决实际视频编辑中的决策层问题,该方法通过结构化补丁和验证器优化视频编辑计划。
📝 Abstract
Practical video editing is not only pixel generation: an editor must turn a brief, a clip pool, music metadata, and hard constraints into an executable timeline. We study this decision layer as \emph{executable video-editing planning} and introduce RefineCut, which, unlike workflow systems that wrap a prompted frontier model, trains a compact open-weight planner for it. The planner edits a typed timeline through structured patches covering clip selection, trimming, ordering, transitions, and duration and music alignment; a deterministic verifier applies each patch and checks it against an explicit constraint ledger. Because editing has no single ground-truth repair, we do not imitate teachers directly: RefineCut replays every multi-teacher branch through the verifier and keeps verifier-best repairs as supervision. A second stage, RefineCut-Evo, lets the student score its own repairs with the verifier and a task rubric and trains on high-margin preference pairs, so the final $8$B planner runs in a closed verifier loop with no teacher calls at inference. On RefineCut-Bench ($3{,}578$ tasks, $7{,}971$ captioned clips, $499$ music tracks, explicit ledgers), verifier-replayed distillation lifts the planner from $0.620$ to $0.858$ on the protocol-specific Video-Editing Score and RefineCut-Evo reaches $0.924$; the gain transfers to Llama-3.1-8B and GLM-4-9B, and in the same closed loop the $8$B planner matches or exceeds its frontier teachers. Code and RefineCut-Bench are publicly released; see the Data Availability statement.
Problem

Research questions and friction points this paper is trying to address.

executable video-editing planning
constraint ledger
verifier-grounded learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

open-weight planner
executable video-editing planning
verifier-grounded learning
structured patches
closed verifier loop
H
Haoyu Wang
School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen; vivo AI Lab
C
Cheng Feng
University of the Chinese Academy of Sciences
L
Liuyang Bian
vivo AI Lab
R
Ruiyang Huang
Southeast University; Peking University
L
Lei Wei
Peking University
Y
Yafei Wen
vivo AI Lab
Xiaoxin Chen
Xiaoxin Chen
Coriell Institute for Medical Research
X
Xiaoying Tang
School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen; Shenzhen Future Network of Intelligence Institute (FNii-Shenzhen); Guangdong Provincial Key Laboratory of Future Networks of Intelligence, CUHK-Shenzhen