Zero-Shot Video Restoration and Enhancement with Text-to-Image Latent Diffusion Models and Multi-Modal References

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种使用文本到图像潜在扩散模型和多模态参考的零样本视频修复与增强框架,解决了视频修复中的时间闪烁问题,并提高了推理速度和时间一致性。
📝 Abstract
Zero-shot image restoration methods with text-to-image latent diffusion models have achieved great success in universal image restoration tasks without training. However, applying them to video restoration will result in severe temporal flickering. In this paper, we propose a novel framework for zero-shot video restoration and enhancement which uses a text-to-image latent diffusion model and multi-modal references. Through the proposed dual prompt tuning inversion and sampling, the inference time can be reduced to nearly 1/3 of the original. The performance and temporal consistency can be also significantly stregthened. By using the proposed texture-aware video token merging, the temporal correlation between frames can be further utilized to improve the temporal consistency. We futher propose the referenced self-attention and referenced token merging to support image reference. Experimental results demonstrate the superiority of the proposed method in restoring and enhancing temporally consistent videos.
Problem

Research questions and friction points this paper is trying to address.

zero-shot video restoration
temporal flickering
text-to-image latent diffusion model
Innovation

Methods, ideas, or system contributions that make the work stand out.

text-to-image latent diffusion model
multi-modal references
dual prompt tuning inversion and sampling
texture-aware video token merging
referenced self-attention
C
Cong Cao
School of Electrical and Information Engineering, Tianjin University, Tianjin, China
H
Huanjing Yue
School of Electrical and Information Engineering, Tianjin University, Tianjin, China
X
Xin Liu
Computer Vision and Pattern Recognition Laboratory, School of Engineering Science, Lappeenranta-Lahti University of Technology LUT, Lappeenranta, Finland
J
Jingyu Yang
School of Electrical and Information Engineering, Tianjin University, Tianjin, China