Preserving Knowledge across Space and Time for Continual Video Deepfake Detection

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种名为MSFD的框架,通过在频域中分解视频特征为时空模态,并采用跨模态解相关损失,解决了连续视频深度伪造检测中的适应性和性能保持问题。
📝 Abstract
The continuous emergence of high-quality video deepfakes requires detectors that continually adapt to new forgery patterns, yet existing approaches, which are designed for deepfake images, fail to capture video-specific cues. Unlike deepfake images that contain only spatial artifacts, deepfake videos leave distinct evidence along both spatial and temporal axes, necessitating the separate preservation of each modality during sequential model updates. To overcome this limitation, we introduce a continual deepfake video detection framework, Modality-Specific Frequency Distillation (MSFD), that explicitly decomposes video features into spatial, temporal, and spatiotemporal modalities in the frequency domain. This decomposition enables independent preservation of each modality, as different deepfake video types exhibit varying reliance on spatial and temporal cues across tasks. Furthermore, MSFD adopts a cross-modality decorrelation loss that encourages spatiotemporal representations to remain orthogonal to single-modality cues. Extensive experiments show that our framework achieves stronger adaptation and preserves performance more effectively than state-of-the-art methods across diverse continual deepfake video scenarios.
Problem

Research questions and friction points this paper is trying to address.

deepfake video detection
continual learning
spatial and temporal cues
Innovation

Methods, ideas, or system contributions that make the work stand out.

Modality-Specific Frequency Distillation
cross-modality decorrelation loss
continual deepfake video detection
🔎 Similar Papers
No similar papers found.