DiffIR2VR-Zero: Zero-Shot Video Restoration with Diffusion-based Image Restoration Models
To address temporal inconsistency arising from direct transfer of pretrained image diffusion inpainting models to video domains, this paper proposes a zero-shot video inpainting framework that reuses arbitrary 2D image diffusion inpainting models without fine-tuning. The method comprises two core innovations: (1) a hierarchical token fusion strategy that enforces inter-frame semantic alignment in the latent space; and (2) a joint optical-flow-guided and feature nearest-neighbor matching mechanism to enhance motion modeling robustness. Crucially, the approach eliminates the need for retraining across diverse degradation types—including 8× super-resolution and Gaussian noise with σ=75—achieving superior performance over fully supervised methods under extreme degradations. Moreover, it demonstrates significant cross-dataset generalization capability, validating its effectiveness beyond domain-specific training.