Beyond Ambiguous Visual Cues: Studying Physiological Disruptions and Cross-Modal Inconsistencies in Deepfake Videos
研究通过构建高保真deepfake视频,分析生理信号和面部行为的不一致性,并提出一种双向共注意融合检测器来提高deepfake检测准确率。
研究通过构建高保真deepfake视频,分析生理信号和面部行为的不一致性,并提出一种双向共注意融合检测器来提高deepfake检测准确率。
本文提出一种通过稀疏竞争促进神经元群组间的竞争来诱导模块化结构的方法,以提高深度神经网络的可解释性和训练效率,无需模块级监督即可在ImageNet-100和CIFAR-100上形成专门化的模块。
该研究解决了关于二次APN函数的猜想,并通过引入与扭曲函数相关的新二次(n,n)函数,探讨了其代数度及相关性质。
This work addresses the challenge of deploying vision Transformer-based facial emotion recognition models on edge devices, where their O(N²) computational complexity poses significant limitations. To this end, we propose Sparse Attention to Emotion (SAE), the first approach to incorporate sparse attention mechanisms into this task. SAE dynamically prunes image tokens irrelevant to emotion discrimination, retaining only those from critical regions such as the eyes and mouth. Experimental results demonstrate that SAE achieves a new state-of-the-art accuracy on the RAF-DB dataset using approximately 10% of the original tokens, while reducing computational complexity by up to 90%. This substantial efficiency gain markedly enhances the model’s practicality for real-world deployment without compromising performance.
This work addresses the challenging task of recognizing individual contradiction and hesitation states in videos—a key problem in affective computing and human-computer interaction—by proposing a novel multimodal approach that integrates textual, audio, and visual modalities. The method leverages F2LLM-v2-0.6B, WavLM-Large, and VideoMAE V2 to extract modality-specific features and introduces a fusion architecture combining bidirectional cross-attention with gated multimodal units to effectively model complementary cross-modal information. Evaluated on the ABAW11 challenge, the proposed approach achieves a Macro F1 score of 0.7394 on the validation set, representing an 11.0% improvement over the best single-modality baseline and significantly outperforming both unimodal and zero-shot methods.
研究通过构建高保真deepfake视频,分析生理信号和面部行为的不一致性,并提出一种双向共注意融合检测器来提高deepfake检测准确率。
本文提出一种通过稀疏竞争促进神经元群组间的竞争来诱导模块化结构的方法,以提高深度神经网络的可解释性和训练效率,无需模块级监督即可在ImageNet-100和CIFAR-100上形成专门化的模块。
该研究解决了关于二次APN函数的猜想,并通过引入与扭曲函数相关的新二次(n,n)函数,探讨了其代数度及相关性质。
This work addresses the challenge of deploying vision Transformer-based facial emotion recognition models on edge devices, where their O(N²) computational complexity poses significant limitations. To this end, we propose Sparse Attention to Emotion (SAE), the first approach to incorporate sparse attention mechanisms into this task. SAE dynamically prunes image tokens irrelevant to emotion discrimination, retaining only those from critical regions such as the eyes and mouth. Experimental results demonstrate that SAE achieves a new state-of-the-art accuracy on the RAF-DB dataset using approximately 10% of the original tokens, while reducing computational complexity by up to 90%. This substantial efficiency gain markedly enhances the model’s practicality for real-world deployment without compromising performance.
This work addresses the challenging task of recognizing individual contradiction and hesitation states in videos—a key problem in affective computing and human-computer interaction—by proposing a novel multimodal approach that integrates textual, audio, and visual modalities. The method leverages F2LLM-v2-0.6B, WavLM-Large, and VideoMAE V2 to extract modality-specific features and introduces a fusion architecture combining bidirectional cross-attention with gated multimodal units to effectively model complementary cross-modal information. Evaluated on the ABAW11 challenge, the proposed approach achieves a Macro F1 score of 0.7394 on the validation set, representing an 11.0% improvement over the best single-modality baseline and significantly outperforming both unimodal and zero-shot methods.