EditaLive! Unified Character Video Editing for Live Streaming

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决实时直播中人物视频编辑问题,提出EditaLive框架,通过改进预训练模型和设计自卷积蒸馏策略实现低延迟实时编辑。
📝 Abstract
Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically depend on multiple offline inference steps, making them unsuitable for real-time interaction. We propose EditaLive, a novel framework for real-time streaming character video editing. In detail, we start from a pretrained image animation model (Wan-Animate), which naturally decouples appearance from motion, and repurpose it as the base model for instruction-based human-centric video editing by reference frame editing and video reconstruction via the collected CharEdit-50K dataset. Besides, we adapt the model from offline bidirectional to causal streaming generation, and design an aligned self-rollout distillation strategy that compresses the model into a two-step sampler, where fixed RoPE and align forcing reduce training--inference discrepancies, and first-frame preserved sparse attention filters redundant historical information to mitigate appearance drift. Extensive experiments demonstrate that EditaLive delivers state-of-the-art editing performance with faithful preservation of facial expressions and low-latency real-time streaming inference.
Problem

Research questions and friction points this paper is trying to address.

live streaming
video editing
facial-expression inconsistencies
real-time interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

real-time streaming
human-centric video editing
causal streaming generation
aligned self-rollout distillation
sparse attention
🔎 Similar Papers
No similar papers found.