Face-voice Association across LAnguages and Gender (FLAG) 2027 Challenge Evaluation Plan
研究通过FLAG 2027挑战赛评估面部-声音关联模型在跨语言和性别条件下的表现,旨在促进能超越语言和性别限制识别说话人身份的模型发展。
研究通过FLAG 2027挑战赛评估面部-声音关联模型在跨语言和性别条件下的表现,旨在促进能超越语言和性别限制识别说话人身份的模型发展。
This study investigates whether popularity calibration genuinely enhances user experience in music recommendation and examines its reliability across varying levels of user listening history and item familiarity. The authors construct three types of playlists—high-popularity, low-popularity, and calibrated—and employ a controlled naive recommender to generate personalized lists. Calibration is quantified using Jensen–Shannon divergence (JSD), and subjective user feedback is collected through controlled experiments. This work presents the first systematic validation of JSD’s stability with respect to real users’ perceived calibration. Results indicate that while users can discern differences in popularity, they do not exhibit a significant preference for calibrated recommendations. Moreover, computed popularity labels show only weak alignment with users’ subjective judgments, and the relationship between JSD and perceived calibration is significantly moderated by item familiarity, playlist composition, and the availability of historical interaction data.
研究通过FLAG 2027挑战赛评估面部-声音关联模型在跨语言和性别条件下的表现,旨在促进能超越语言和性别限制识别说话人身份的模型发展。
This study investigates whether popularity calibration genuinely enhances user experience in music recommendation and examines its reliability across varying levels of user listening history and item familiarity. The authors construct three types of playlists—high-popularity, low-popularity, and calibrated—and employ a controlled naive recommender to generate personalized lists. Calibration is quantified using Jensen–Shannon divergence (JSD), and subjective user feedback is collected through controlled experiments. This work presents the first systematic validation of JSD’s stability with respect to real users’ perceived calibration. Results indicate that while users can discern differences in popularity, they do not exhibit a significant preference for calibrated recommendations. Moreover, computed popularity labels show only weak alignment with users’ subjective judgments, and the relationship between JSD and perceived calibration is significantly moderated by item familiarity, playlist composition, and the availability of historical interaction data.