Real-Time Video Prediction With Fast Video Interpolation Model and Prediction Training
To address perceptual latency degradation in real-time video transmission, which impairs interactive user experience, this paper proposes IFRVP—a zero-latency video prediction framework. Methodologically, we design IFRNet, a lightweight convolutional architecture incorporating ELAN-based residual modules to balance accuracy and efficiency, and introduce three novel frame interpolation training paradigms specifically tailored for predictive tasks. Furthermore, we propose a mid-level feature refinement mechanism to enable end-to-end inter-frame interpolation modeling. Experimental results demonstrate that IFRVP achieves a state-of-the-art trade-off between prediction accuracy and inference speed, enabling real-time prediction at over 30 FPS and significantly reducing end-to-end perceptual latency. The source code and demonstration videos are publicly available.