MDN-Control: Mask-Depth-Noise Guided Region Control for Multi-Subject Video Editing

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多主体视频编辑中的属性泄露和遮挡模糊问题,提出MDN-Control框架,通过掩码引导定位、深度感知遮挡控制及噪声潜提示来改善目标定位、几何遮挡处理和外观初始化。
📝 Abstract
Multi subject video editing modifies designated subjects while preserving non target content, but faces cross subject attribute leakage, and occlusion ambiguity. Existing approaches rely on masks and struggle to distinguish overlapping subjects or ensure consistent generation. To address these limitations, we propose MDN-Control, a training free framework jointly controlling target localization, occlusion geometry, and appearance initialization. Specifically, mask-guided localization provides consistent target localization, while depth-aware occlusion control resolves ambiguous boundaries between overlapping subjects. We further introduce noise latent prompting, which retrieves Gaussian initializations from a noise library for prompt relevant priors. Experiments on MSVBench show that MDN-Control achieves the lowest CM-Err and the highest Q-Edit, while maintaining competitive text alignment and temporal consistency, demonstrating the effectiveness of combining spatial, geometric, and latent priors for multi subject video editing.
Problem

Research questions and friction points this paper is trying to address.

multi-subject video editing
cross-subject attribute leakage
occlusion ambiguity
Innovation

Methods, ideas, or system contributions that make the work stand out.

mask-guided localization
depth-aware occlusion control
noise latent prompting
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
J
Jiayi Yu
School of Cyber Science and Engineering, Wuhan University
X
Xi Ye
School of Cyber Science and Engineering, Wuhan University
Lina Wang
Lina Wang
Professor, Wuhan University
Computer Security
Y
Yunkun Xia
School of Cyber Science and Engineering, Wuhan University