An Inclusive and Lightweight Approach to Federated Continual Learning for Cultural Heritage
本文提出了一种轻量级的联邦持续学习方法FedCurv-DR,旨在解决文化遗产数据分布广泛、受限且不断变化的问题,通过累积参数重要性估计来保护已学知识并减少通信和计算开销。
本文提出了一种轻量级的联邦持续学习方法FedCurv-DR,旨在解决文化遗产数据分布广泛、受限且不断变化的问题,通过累积参数重要性估计来保护已学知识并减少通信和计算开销。
This study addresses the limited robustness of large vision-language models against spurious correlations—such as single- and multi-attribute biases—and the unclear relationship between model scale and debiasing efficacy. Through a systematic empirical analysis of 194 publicly available models, the authors evaluate how model size, training data, and architectural design influence bias sensitivity on ImageNet, CelebA, and UrbanCars benchmarks. They find that increasing model scale has nearly no effect on mitigating complex multi-attribute biases (ρ = 0.05), whereas high-quality, large-scale training data consistently improves worst-group accuracy by up to 25%. The impact of architectural choices, however, is highly dependent on the specific bias type and its spatial distribution.
Current automated facial age verification systems are vulnerable to circumvention by minors using simple appearance modifications, such as drawing beards or applying lipstick. This study presents the first systematic evaluation of the robustness of seven state-of-the-art visual and multimodal models against four types of minimalistic appearance manipulations across three datasets, employing lightweight linear probes to mitigate dataset bias. Experimental results demonstrate that such manipulations can cause up to 61% of genuine negative samples to be misclassified as positive. Moreover, vulnerability exhibits significant demographic disparities: individuals of Indian descent are more susceptible to beard-based spoofing, and females experience higher overall false acceptance rates than males. This work quantifies, for the first time, the performance degradation of age verification models under minimal adversarial appearance changes and reveals socially consequential disparities in their failure modes.
This work addresses the limited interpretability of current deepfake detectors, which often function as black-box models and thus lack the transparency required for high-stakes applications such as legal proceedings. To bridge this gap, the authors propose a novel post-hoc explainable artificial intelligence (XAI) approach based on Encoder-Decoder Direction Pairs (EDDP), which globally disentangles the semantic concepts learned by the detector and their underlying read-write mechanisms without requiring manual annotations. This method uniquely enables spatially aware concept localization and counterfactual analysis, effectively uncovering the fine-grained authentic and forged features implicitly captured by the model. As a result, it substantially enhances both the interpretability and decision transparency of deepfake detection systems.
This work addresses the task of generating semantically consistent images from news headlines by proposing ACIG, a test-time, model-agnostic Actor-Critic image generation method. ACIG introduces the Actor-Critic paradigm from reinforcement learning into text-to-image synthesis for the first time, iteratively refining textual prompts through a self-feedback loop: the Actor proposes prompt modifications, while the Critic evaluates the semantic alignment and quality of the resulting images and provides feedback to guide subsequent refinements. Notably, this optimization occurs entirely at inference time without requiring additional training. ACIG is flexibly compatible with existing text-to-image models and achieved top performance on the official leaderboard of the MediaEval NewsImages 2026 challenge, significantly enhancing the semantic consistency between generated images and their corresponding news content.
本文提出了一种轻量级的联邦持续学习方法FedCurv-DR,旨在解决文化遗产数据分布广泛、受限且不断变化的问题,通过累积参数重要性估计来保护已学知识并减少通信和计算开销。
This study addresses the limited robustness of large vision-language models against spurious correlations—such as single- and multi-attribute biases—and the unclear relationship between model scale and debiasing efficacy. Through a systematic empirical analysis of 194 publicly available models, the authors evaluate how model size, training data, and architectural design influence bias sensitivity on ImageNet, CelebA, and UrbanCars benchmarks. They find that increasing model scale has nearly no effect on mitigating complex multi-attribute biases (ρ = 0.05), whereas high-quality, large-scale training data consistently improves worst-group accuracy by up to 25%. The impact of architectural choices, however, is highly dependent on the specific bias type and its spatial distribution.
Current automated facial age verification systems are vulnerable to circumvention by minors using simple appearance modifications, such as drawing beards or applying lipstick. This study presents the first systematic evaluation of the robustness of seven state-of-the-art visual and multimodal models against four types of minimalistic appearance manipulations across three datasets, employing lightweight linear probes to mitigate dataset bias. Experimental results demonstrate that such manipulations can cause up to 61% of genuine negative samples to be misclassified as positive. Moreover, vulnerability exhibits significant demographic disparities: individuals of Indian descent are more susceptible to beard-based spoofing, and females experience higher overall false acceptance rates than males. This work quantifies, for the first time, the performance degradation of age verification models under minimal adversarial appearance changes and reveals socially consequential disparities in their failure modes.
This work addresses the limited interpretability of current deepfake detectors, which often function as black-box models and thus lack the transparency required for high-stakes applications such as legal proceedings. To bridge this gap, the authors propose a novel post-hoc explainable artificial intelligence (XAI) approach based on Encoder-Decoder Direction Pairs (EDDP), which globally disentangles the semantic concepts learned by the detector and their underlying read-write mechanisms without requiring manual annotations. This method uniquely enables spatially aware concept localization and counterfactual analysis, effectively uncovering the fine-grained authentic and forged features implicitly captured by the model. As a result, it substantially enhances both the interpretability and decision transparency of deepfake detection systems.
This work addresses the task of generating semantically consistent images from news headlines by proposing ACIG, a test-time, model-agnostic Actor-Critic image generation method. ACIG introduces the Actor-Critic paradigm from reinforcement learning into text-to-image synthesis for the first time, iteratively refining textual prompts through a self-feedback loop: the Actor proposes prompt modifications, while the Critic evaluates the semantic alignment and quality of the resulting images and provides feedback to guide subsequent refinements. Notably, this optimization occurs entirely at inference time without requiring additional training. ACIG is flexibly compatible with existing text-to-image models and achieved top performance on the official leaderboard of the MediaEval NewsImages 2026 challenge, significantly enhancing the semantic consistency between generated images and their corresponding news content.