🤖 AI Summary
This study addresses the perceptual gap wherein textual instructions fail to convey fine-grained textures and complex dynamics by pioneering a visual context editing paradigm. Leveraging a newly constructed dataset of 400,000 samples and a unified framework, we propose modality-adaptive semantic distillation and dual-context injection mechanisms to effectively integrate audio-visual signals for precise editing. The proposed method achieves state-of-the-art performance on both basic instruction-following and visual context editing tasks, significantly enhancing semantic alignment and detail fidelity in video editing. These results validate the effectiveness of the new paradigm and provide critical technical support for multimodal video generation, bridging the limitations of text-only conditioning in complex generative tasks.
📝 Abstract
Despite progress in instruction-based video editing, unimodal textual instructions inherently struggle to convey fine-grained textures and complex dynamics. To bridge this perceptual gap, we propose Visual In-context Editing, a new paradigm elevating video editing from textual instructions to multi-modal visual guidance encompassing single image, image pair, and video pair. To facilitate this paradigm, we curate VicEdit-400K, the first large-scale dataset for visual in-context video editing. We develop an automated pipeline to generate 400K high-quality samples across ten task types, ensuring superior visual fidelity and semantic consistency through multi-dimensional filtering. Leveraging this foundation, we introduce VicEdit, a unified framework to bridge visual and textual contexts. To adaptively extract editing semantics from heterogeneous references, we design Modality-Adaptive Semantic Distillation, which produces modality-specific semantic tokens from visual references. These tokens are then synergistically integrated with textual instructions through Dual-Context Injection, enabling the generation process to benefit from both visual and textual signals. Extensive evaluations on VicEditBench demonstrate that VicEdit achieves state-of-the-art performance across both basic instruction editing and visual in-context editing tasks, establishing visual in-context learning as a powerful and controllable paradigm for video editing.