Benchmarking Cross-Domain Audio-Visual Deception Detection
Current audio-visual spoofing detection methods suffer from poor cross-scenario generalization and lack a standardized cross-domain evaluation benchmark. To address this, we introduce the first unified, standardized cross-domain benchmark for audio-visual spoofing detection, supporting both single-source-to-single-target and multi-source-to-single-target domain adaptation settings. We propose MM-IDGM, a gradient-coordinated optimization algorithm, and Attention-Mixer, a novel multimodal fusion architecture. Additionally, we design three novel multi-source domain sampling strategies and integrate OpenSMILE/ResNet-50 feature extractors with CNN/RNN/Transformer backbones. Extensive experiments demonstrate that our approach achieves an average accuracy improvement of 5.2% under the multi-source-to-single-target setting, significantly enhancing cross-domain generalization. The benchmark and methodology provide a reproducible, comparable, and realistic evaluation framework for practical deployment.