MS-MFAD : Multimodal large language models for Face Anti-spoofing Detection

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决面部生物识别系统面临的复合威胁,提出基于多模态大语言模型的MS-MFAD方法,通过细粒度像素-语义锚定机制提高泛化能力和可审计性。
📝 Abstract
Facial biometric recognition systems currently face compound threats intertwining generative AI and high-fidelity physical spoofing. Existing defenses suffer from systemic bottlenecks, including poor generalization, non-auditable reasoning, and reliance on massive, low-quality datasets. To address these challenges, we propose Multimodal Large Language Models (MFAD) for face anti-spoofing detection, an explainable reasoning system for Unified Face Anti-Spoofing Detection (UFAD), accompanied by a semantic-level annotation benchmark. Unlike methods relying on external tools or coarse alignment, MFAD activates the intrinsic reasoning capabilities of Multimodal Large Language Models (MLLMs) via a fine-grained pixel-semantic anchoring mechanism. This eliminates localization hallucinations and ensures auditable reasoning paths. We introduce a cross-attack semantic-level unified annotation paradigm: by annotating only 1,000 precise masks per attack category, we generate reasoning evidence chains strictly corresponding to spoofed regions. Supervised fine-tuning on the Qwen-VL foundation model demonstrates that, using limited high-quality samples, the system achieves a 40-50% relative reduction in in-domain ACER and restricts cross-domain performance degradation to within 11.62%/5.23%, significantly outperforming existing frameworks. Furthermore, under white-box adversarial attacks, detection accuracy drops by only 3.2%, validating the robustness of semantic anchoring compared to models trained on massive short-text data. Domain practitioners rated the evidence reliability of reasoning paths at 4.57/5, with inference latency satisfying real-time deployment requirements. These results confirm that a few-shot, high-quality semantic annotation paradigm is effective for building trustworthy, explainable, and cost-efficient UFAD systems.
Problem

Research questions and friction points this paper is trying to address.

Facial biometric recognition
Generative AI
Physical spoofing
Systemic bottlenecks
Multimodal Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Large Language Models
pixel-semantic anchoring mechanism
cross-attack semantic-level unified annotation paradigm
fine-tuning
white-box adversarial attacks
💼 Related Jobs
No related jobs found.
X
Xiaoyong Yu
Mashang Consumer Finance Co., Ltd., 4th to 8th Floors, Building B2, Yuxing Plaza, No. 52 Huangshan Avenue Middle Section, Chongqing, 401121, Chongqing, China
R
Rongzhen Li
Mashang Consumer Finance Co., Ltd., 4th to 8th Floors, Building B2, Yuxing Plaza, No. 52 Huangshan Avenue Middle Section, Chongqing, 401121, Chongqing, China
Shuming Shi
Shuming Shi
Tencent AI Lab
NLPtext understandingknowledge miningtext generationweb search
Xinge You
Xinge You
Professor of School of Electronics Information and Communications, Huazhong University of Science
Computer VisionPattern RecognitionMachine LearningWavelet Analysis and its Applications