Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决深伪图像检测泛化问题,提出UCF-Net,融合CLIP和DINO的特征,并基于不确定性进行加权融合。
📝 Abstract
The growing realism and accessibility of manipulated and generated faces threaten the trustworthiness of digital media. To detect such forgeries, deepfake detectors based on vision foundation models have shown promising performance, but they typically rely on a single pretrained representation and are prone to overfitting to particular training distributions. To improve generalization to unseen forgeries, we propose UCF-Net, an uncertainty-aware cascaded fusion network that harnesses CLIP's language-aligned semantic priors and DINO's self-supervised visual-structure priors. UCF-Net extracts hierarchical features across Transformer depths, uses layer-wise expert aggregation to adaptively combine each encoder's multi-level cues, and performs weighted fusion of the resulting representations based on entropy-derived uncertainty. We further consolidate public deepfake datasets into a unified benchmark of approximately 4M images and construct a separate cross-generator evaluation set with over 8K face images from eight recent generators. On the unified benchmark, UCF-Net achieves the best mean AUC among the evaluated methods in both in-domain and cross-domain evaluations. On the cross-generator set, it adapts effectively with limited target-domain data, although zero-shot transfer remains challenging.
Problem

Research questions and friction points this paper is trying to address.

deepfake detection
generalization
pretrained representation
overfitting
unseen forgeries
Innovation

Methods, ideas, or system contributions that make the work stand out.

uncertainty-aware
cascaded fusion network
CLIP
DINO
generalizable deepfake detection
X
Xuechao Zou
Beijing Jiaotong University
Y
Yi Zhou
Beijing Jiaotong University
K
Kai Li
Tsinghua University
S
Shun Zhang
Beijing Jiaotong University
Y
Yuhui Chen
Ant Group
Congyan Lang
Congyan Lang
Beijing Jiaotong University
computer vision
J
Junliang Xing
Tsinghua University