Vision Language Model Fusion for Explainable Face Recognition

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了通过融合多个视觉-语言模型来提高面部识别准确性并提供更丰富解释的方法,以支持透明度、可审计性和错误分析。
📝 Abstract
Responsible deployment of face verification systems requires more than accurate decisions: systems should also provide interpretable and auditable evidence that enables users to understand, assess, and challenge their decisions. Vision-language models (VLMs) provide a promising foundation for explainable face recognition by combining visual analysis with natural-language reasoning. However, relying on a single model may further limit the decision accuracy as well as provided explanations. This work therefore investigates whether multiple VLMs can be combined to improve recognition accuracy, and to enrich the explanations associated with those decisions. This work evaluates four VLMs as standalone face verification systems and subsequently proposes a fusion framework, where two source models provide similarity scores and textual justifications and a third VLM acts as a decider model. Four different fusion scenarios are considered, progressively providing the decider model with scores, justifications, face images, and combinations of these modalities. Overall, the findings suggest that the value of multi-VLM fusion extends beyond recognition performance. VLMs can provide complementary justifications and perspectives that enable richer explanations of face recognition decisions, supporting greater transparency, auditability, and error analysis. This is relevant to the development of responsible explainable face verification systems, where users and operators should be able to understand not only the final decision but also the evidence and potential sources underlying it. The proposed multimodal VLM, which combines decision scores, explanations, and face images, achieves higher recognition accuracy than state-of-the-art VLMs and domain-specific face recognition models, while also providing fused explanations that are expected to be more robust than those generated by individual VLMs.
Problem

Research questions and friction points this paper is trying to address.

Explainable Face Recognition
Vision-language Models
Decision Accuracy
Interpretable Evidence
Transparency
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal VLM
face recognition
explainable AI
decision fusion
transparency
🔎 Similar Papers
No similar papers found.
A
Ana Estrada-Real
da/sec Biometrics and Security Research Group, Hochschule Darmstadt, Germany
L
Lydia Alapatt
da/sec Biometrics and Security Research Group, Hochschule Darmstadt, Germany
Christoph Busch
Christoph Busch
Professor for Biometrics, Norwegian University of Science and Technology (NTNU)
Biometrics
Christian Rathgeb
Christian Rathgeb
Hochschule Darmstadt
Pattern RecognitionMachine LearningBiometrics