🤖 AI Summary
This work addresses the vulnerability of vision-language and generative image models to geometric transformations such as input rotation, which can degrade robustness and exacerbate demographic biases. The study presents the first systematic investigation into how rotation perturbations jointly impact fairness and robustness in multimodal models. To mitigate these issues, the authors propose an integrated framework combining data augmentation, representation alignment, and model regularization. Evaluated across multiple benchmark datasets, the approach substantially enhances robustness to rotation, effectively curbs bias amplification, and maintains or even improves overall task performance, thereby achieving a synergistic optimization of robustness, fairness, and accuracy.
📝 Abstract
Vision-Language Models (VLMs) and generative image models have achieved remarkable performance across multimodal tasks, yet their robustness and fairness under input transformations remain insufficiently explored. This work investigates bias propagation and robustness degradation in state-of-the-art vision-language and generative models, with a particular focus on image rotation and distributional shifts. We analyze how rotation-induced perturbations affect model predictions, confidence calibration, and demographic bias patterns. To address these issues, we propose rotation-robust mitigation strategies that combine data augmentation, representation alignment, and model-level regularization. Experimental results across multiple datasets demonstrate that the proposed methods significantly improve robustness while reducing bias amplification without sacrificing overall performance. This study highlights critical limitations of current multimodal systems and provides practical mitigation techniques for building more reliable and fair AI models.