Adversarially Robust CLIP Models Can Induce Better (Robust) Perceptual Metrics
Neural perceptual metrics like CLIP exhibit insufficient robustness against adversarial attacks, limiting their reliability in safety-critical zero-shot evaluation tasks. Method: We propose R-CLIP$_ extrm{F}$, an unsupervised adversarial fine-tuning framework that jointly optimizes CLIP’s feature space for robustness and performs self-supervised adaptation under adversarial perturbations—without requiring labeled data. It further introduces feature and text inversion mechanisms to enhance interpretability and enable visual concept visualization. Contribution/Results: R-CLIP$_ extrm{F}$ is the first method achieving both high robustness and high discriminability in zero-shot perceptual similarity modeling. Experiments demonstrate its superior performance over state-of-the-art metrics on zero-shot perceptual assessment, robust vision–language retrieval, and NSFW content detection. Crucially, it maintains high accuracy under adversarial perturbations while preserving original-image performance—bridging the robustness–accuracy trade-off without supervision.