SNM-VFI: Symmetric Nonlinear Motion-Guided Generative Video Frame Interpolation

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of poor motion controllability, low perceptual quality, and temporal inconsistency in video frame interpolation by proposing a training-free interpolation framework. It leverages a pretrained optical flow model to construct symmetric nonlinear motion-guided frames, which serve as latent-space priors to iteratively steer a pretrained video diffusion model for high-fidelity and motion-coherent intermediate frame synthesis. The method innovatively integrates symmetric nonlinear motion modeling with a pretrained video diffusion model and introduces a confidence map fusion mechanism that balances structural reliability and textural realism in ambiguous regions such as occlusions and object boundaries. Extensive experiments on standard benchmarks—including DAVIS, Sintel, and KITTI—demonstrate superior performance in perceptual quality, reconstruction accuracy, and temporal consistency.
📝 Abstract
We propose Symmetric Nonlinear Motion-guided Generative Video Frame Interpolation (SNM-VFI), a training-free framework for motion-controllable generative video frame interpolation with pre-trained optical flow and video diffusion models. Unlike conventional diffusion-based VFI methods that synthesize intermediate frames from random noise, SNM-VFI guides the generative process with correspondence-aware frames produced by a symmetric nonlinear motion model. Specifically, we first utilize a pre-trained optical flow model to construct multi-frame nonlinear flow-based intermediate frames and confidence maps. These flow-guided frames are then encoded as latent priors to initialize and iteratively guide a pre-trained Video Diffusion model, enabling the diffusion model to preserve dense motion correspondence while improving perceptual realism. To further enhance output quality, we employ confidence maps to fuse structurally reliable flow-based predictions with diffusion-generated details in uncertain regions such as occlusions and object boundaries. Extensive evaluations on challenging benchmarks, including DAVIS, Sintel, and KITTI, demonstrate that SNM-VFI achieves strong perceptual quality, competitive reconstruction accuracy, and robust temporal coherence across diverse motion scenarios.
Problem

Research questions and friction points this paper is trying to address.

Video Frame Interpolation
Motion Correspondence
Diffusion Models
Optical Flow
Temporal Coherence
Innovation

Methods, ideas, or system contributions that make the work stand out.

video frame interpolation
diffusion model
optical flow
motion guidance
training-free