Over++: Generative Video Compositing for Layer Interaction Effects
In professional video compositing, environment interactions—such as shadows, reflections, dust, and splashes—between foreground and background have traditionally relied on labor-intensive manual creation. Existing video generation models struggle to inject photorealistic interactions while preserving the input video; conversely, video inpainting methods suffer from either requiring frame-wise manual masks or producing geometrically distorted outputs. To address this, we introduce “enhanced compositing” as a novel task, construct the first paired dataset of videos with environment interaction effects, and propose a self-supervised, unpaired training strategy jointly guided by text prompts, segmentation masks, and keyframes. Our method integrates video diffusion models, unsupervised spatiotemporal consistency modeling, lightweight mask fusion, and keyframe-guided distillation. Experiments demonstrate that our approach generates diverse, high-fidelity semi-transparent interactions under data constraints, achieving state-of-the-art performance in both interaction realism and source scene fidelity.