Bridging Event Streams and DiT: Event-Guided Video Frame Interpolation

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of motion blur, structural distortion, and temporal inconsistency in video frame interpolation under large temporal gaps and complex motion. The authors propose a novel approach that leverages pre-trained image-to-video diffusion models without requiring retraining. By incorporating high-temporal-resolution motion cues from event cameras, they design a lightweight adapter architecture that fuses Image Warped Events (IWEs) with bidirectional sparse optical flow to generate spatiotemporally aligned structural and motion guidance signals, which are injected into the latent diffusion model. This method represents the first effective integration of event streams with pre-trained DiT-based video generation models, achieving state-of-the-art performance on both real-world and synthetic benchmarks, significantly improving interpolation fidelity and temporal coherence.
📝 Abstract
Latent diffusion models have recently advanced video frame interpolation by synthesizing intermediate frames between input images. However, handling large temporal gaps and complex motion remains challenging, often resulting in motion blur, structural distortions, and temporal inconsistencies. Event cameras provide high-temporal-resolution motion cues that are well suited for bridging these gaps and improving interpolation quality. To exploit this advantage without training an event-assisted model from scratch, we propose an adapter-based framework that incorporates event-derived cues into a pre-trained image-to-video diffusion model with minimal architectural changes. Specifically, our method leverages Image Warped Events (IWEs) and bidirectional sparse optical flow to provide spatially and temporally aligned guidance during generation. By injecting these event-guided structural and motion cues into the diffusion process, our approach reduces interpolation artifacts and improves both reconstruction fidelity and temporal coherence. Experimental results on real and synthetic benchmarks show that our method consistently outperforms existing state-of-the-art approaches. The project page is at https://joseph-lin-tech.github.io/BridgeEventDiT-VFI/.
Problem

Research questions and friction points this paper is trying to address.

video frame interpolation
temporal gaps
motion blur
structural distortions
temporal inconsistencies
Innovation

Methods, ideas, or system contributions that make the work stand out.

event camera
video frame interpolation
latent diffusion model
adapter-based framework
temporal coherence
🔎 Similar Papers
No similar papers found.