Fair on the Surface? Benchmarking Hidden-Output Fairness Gaps in LLM Recommenders

📅 2026-08-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited scope of fairness evaluation in existing large language model–based recommender systems, which typically focuses only on output-level metrics while overlooking implicit biases embedded in hidden representations. To bridge this gap, we propose FairGap, a novel benchmark that simultaneously assesses output-based fairness (OBS) and internal representation fairness (IBS) through controlled counterfactual identity probes, and introduces a Representation–Output Alignment (ROA) metric for joint diagnosis. Our analysis reveals a pervasive misalignment between OBS and IBS—ROA scores are consistently low (≤0.22)—with some users exhibiting stable outputs despite significant internal representation shifts. Further intervention experiments demonstrate that reducing IBS by up to 8× can paradoxically degrade OBS, highlighting a fundamental tension between fairness at different model layers. The proposed four-quadrant diagnostic framework effectively identifies users with hidden–output mismatches, uncovering latent bias patterns invisible to conventional evaluation approaches.
📝 Abstract
Fairness audits for LLM-based recommenders have largely focused on observable outputs, implicitly assuming that stable recommendations reflect stable internal processing. We challenge this assumption with FairGap, the first benchmark to jointly evaluate recommendation fairness at two levels: observable output shift (OBS) and hidden representation shift (IBS), measured through controlled counterfactual identity probes across gender, age, and race. Their relationship is summarized via Representation-Output Alignment (ROA), with quadrant diagnostics for identifying user-level hidden-output mismatch. Applied to six open-weight LLM families across three domains, FairGap reveals pervasive hidden-output decoupling: ROA rarely exceeds 0.22, and a non-negligible user population shows stable outputs despite substantial internal shifts, a mode that output-only audits cannot detect by design. Further, activation steering that reduces IBS by up to 8x simultaneously worsens OBS, demonstrating a fundamental tension between internal and output-level fairness that existing frameworks are unequipped to diagnose.
Problem

Research questions and friction points this paper is trying to address.

fairness
LLM recommenders
hidden representation
output shift
counterfactual probing
Innovation

Methods, ideas, or system contributions that make the work stand out.

hidden-output fairness
representation-output alignment
counterfactual identity probes
activation steering
fairness benchmarking
🔎 Similar Papers
No similar papers found.