DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过集成多种预训练模型提取多模态特征,并使用混合专家模型解决跨方法泛化的深度伪造检测问题。
📝 Abstract
Audio-visual deepfake detection is an actively studied topic, where one of the main challenges is to develop detectors able to generalize across deepfake generation methods. We conjecture that overfitting can be mitigated by extracting multiple high-level cues from the available audio and visual modalities via pre-trained models. We therefore assemble a wide variety of pre-trained models to extract features that encode mouth movements, face parsing, facial expressions, head pose, gaze tracking, heart rate, audio emotion and speech activity. We further integrate both unimodal and multimodal cues via a Mixture-of-Experts (MoE) backbone to detect deepfakes. We perform in-domain and cross-domain experiments on five benchmarks for deepfake detection (MAVOS-DD, AVLips, PolyGlotFake, BioDeepAV, FakeAVCeleb) to compare our framework (DF-MoE) with state-of-the-art methods. Our results indicate that DF-MoE obtains superior deepfake detection results, surpassing all competing methods. We release our code at https://github.com/vladhondru25/DF-MoE.
Problem

Research questions and friction points this paper is trying to address.

deepfake detection
generalization
multimodal
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Multimodal Features
Deepfake Detection
Generalization
Pre-trained Models
🔎 Similar Papers
No similar papers found.