RA-ClipScore: Making Generative Model Evaluation More Interpretable
This work addresses the limited interpretability of existing evaluation metrics for generative models, which struggle to diagnose generation biases at both semantic attribute and spatial distribution levels. To this end, the paper proposes RA-CLIPScore, the first CLIP-based evaluation framework extended to incorporate spatial alignment. RA-CLIPScore employs a dual-prompt mechanism to disentangle competing attributes and leverages local image patch tokens to model region-wise semantic alignment. By introducing a region-specific single-attribute divergence metric, the method substantially enhances interpretability and aligns more closely with human perception. Experiments demonstrate that RA-CLIPScore exhibits greater robustness under distribution shifts or when textual prompts contain partially irrelevant attributes, and its scores show strong correlation with human judgments of visual diversity.