Mind the Rift: Cross-Scale Coupling Mismatch for AI-Generated Video Detection

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对AI生成视频的检测问题,提出了一种名为RIFT的新方法,通过分析宏观和微观尺度的信息及其跨尺度关系来识别视频真伪。
📝 Abstract
As AI video generators achieve cinematic realism, reliable detection becomes essential for safeguarding digital trust. We identify cross-scale coupling mismatch as a new forensic signal, where scale refers to the level of abstraction (semantic dynamics vs. pixel-level residuals): in natural videos, macro-level temporal dynamics and micro-level residual patterns are intrinsically coupled by the unified imaging physics pipeline, whereas AI generators, whose training objectives do not explicitly preserve this joint distribution, systematically violate this coupling. Detecting such mismatch is challenging because it requires independently extracting information at both scales while simultaneously quantifying their cross-scale relationship. We propose RIFT (Representation Inconsistency Forensics on Trajectories), an orthogonal forensic framework that addresses this through three interlocking components: a macro stream that builds a dynamic baseline of expected temporal evolution via differential geometry and persistent homology on learned manifold trajectories, a micro stream that acts as a sensitive forensic probe via steganalytic filtering and temporal modeling, and a coupling divergence module that measures the conditional dependency between the two streams. Gram-Schmidt orthogonality guarantees the information-theoretic validity of this measurement. Experiments on two benchmarks (VidProM, 120K videos, 7 generators; GenVidBench, 68K videos, 4 generators) demonstrate that RIFT achieves 99.33% and 99.72% F1-score respectively, with 97.87% unseen-generator detection rate in leave-one-out evaluation, while exhibiting encoder agnosticism: scaling from ViT-S/14 (22M) to ViT-L/14 (300M) changes F1 by less than 0.1%, and switching to a different encoder family (DINOv1) reduces F1 by only 0.73 pp. Code is available at https://github.com/Litsay/RIFT
Problem

Research questions and friction points this paper is trying to address.

Cross-Scale Coupling Mismatch
AI-Generated Video Detection
Temporal Dynamics
Residual Patterns
Digital Trust
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-scale coupling mismatch
RIFT
differential geometry
persistent homology
steganalytic filtering
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Siyu Li
Siyu Li
University of Illinois at Chicago
RoboticsMicro-robot swarmsHuman-robot InteractionControl and Motion Planning
J
Jin Yang
School of Cyber Science and Engineering, Sichuan University; School of Information Science and Technology, Xizang University
W
Weiheng Liang
School of Cyber Science and Engineering, Sichuan University