Face-voice Association across LAnguages and Gender (FLAG) 2027 Challenge Evaluation Plan

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过FLAG 2027挑战赛评估面部-声音关联模型在跨语言和性别条件下的表现,旨在促进能超越语言和性别限制识别说话人身份的模型发展。
📝 Abstract
Face--voice association models may rely on language or gender cues in the voice rather than on speaker-specific voice characteristics, which can lead to a performance deterioration when the model has to identify a multilingual speaker or distinguis same-gender speakers. To investigate these issues, we introduce the Face-voice Association across LAnguages and Gender (FLAG) 2027 Challenge. The challenge formulates face--voice association as a cross-modal verification task: given a voice, identify the speaker's face from a ``gallery'' of faces consisting of the speaker's face and a set of negative samples. Models are evaluated on identities not present in the training data (``unseen'') and both for languages present or absent from the training data (``heard'' and ``unheard''). Two evaluation settings are used to test models' reliance on gender: a standard, unconstrained and a gender-constrained one, where the latter uses a same-gender gallery. The performance of existing, baseline models in these settings reveals that models performance degrades under language shifts and in gender-constrained settings, highlighting the need to foster the development of models that capture identity-specific aspects beyond language and gender. The challenge provides a benchmark dataset, pretrained baseline models, and an evaluation framework to advance face--voice association.
Problem

Research questions and friction points this paper is trying to address.

face-voice association
language cues
gender cues
speaker-specific voice characteristics
multilingual speaker
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-modal verification
multilingual speaker identification
gender-constrained evaluation
🔎 Similar Papers
No similar papers found.