Mechanism-Level Evaluation for Vision-Language Models: Controlled Activation-Replacement Diagnosis of Gender Bias

πŸ“… 2026-09-15
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
ζœ¬ζ–‡ι€šθΏ‡ε› ζžœδΈ­δ»‹εˆ†ζžζ–Ήζ³•οΌŒθ―Šζ–­θ§†θ§‰-θ―­θ¨€ζ¨‘εž‹δΈ­ηš„ζ€§εˆ«εθ§ζœΊεˆΆοΌŒζ­η€ΊδΊ†ε„ε±‚ζΏ€ζ΄»ε―ΉεΉ²ι’„ηš„ζ•ζ„Ÿζ€§γ€‚
πŸ“ Abstract
Behavioral benchmarking reveals \emph{what} biases exist in vision-language models but not \emph{which internal components} are most sensitive to targeted intervention, precluding principled intervention. We argue for mechanism-level evaluation as a necessary complement, demonstrating causal mediation analysis as a diagnostic instrument for gender bias. We decompose gender-cue effects into controlled indirect effects attributable to specific-layer activations and direct effects through all other pathways, producing layer-by-layer mechanistic signatures. Across six models spanning three architectural families (LLaVA-1.5, LLaVA-NeXT, InstructBLIP at 7B/13B) and two 8B-scale architectures, three findings emerge: language-layer activations exhibit the greatest output sensitivity under controlled intervention, with the direct component often carrying the opposite sign; architectural choices redistribute layer-wise sensitivity to activation replacement; and counterfactual scores diverge from surface-level scores, exposing implicit associations. Systematic ablation validates internal consistency. An intervention experiment finds that the average indirect effect (AIE) and downstream intervention effectiveness are only weakly correlated (Pearson $r = 0.33$), and the layer with the second-largest AIE produces near-zero bias change---indicating that mechanistic diagnosis captures activation-replacement sensitivity but does not, by itself, identify optimal intervention targets. These results show mechanism-level evaluation captures architecture-specific sensitivity patterns that behavioral benchmarks cannot; pairing both should become standard NLP practice. Code: https://github.com/zhaozhipeng1997/CARD-GenderBias.
Problem

Research questions and friction points this paper is trying to address.

gender bias
mechanism-level evaluation
vision-language models
controlled intervention
Innovation

Methods, ideas, or system contributions that make the work stand out.

mechanism-level evaluation
controlled activation-replacement diagnosis
causal mediation analysis
gender bias
πŸ”Ž Similar Papers
2024-07-03Conference on Empirical Methods in Natural Language ProcessingCitations: 1
Z
Zhipeng Zhao
Ocean University of China
W
Wenxu Wang
Ocean University of China
P
Peishun Liu
Ocean University of China
R
Ruichun Tang
Ocean University of China