🤖 AI Summary
This study investigates how task-oriented "personas" influence the relevance judgments of large language models (LLMs) when employed as evaluators. Leveraging PersonaHub and NVIDIA Nemotron-Personas-USA, the authors construct five distinct evaluator personas and conduct experiments across six LLM backbones on the TREC DL20 and RAG24 datasets, using UMBRELA as the baseline. Through local rank displacement and system-level consistency analyses, the work pioneers the use of personas as controllable probes to reveal that LLM evaluation exhibits structured sensitivity rather than uniform bias: high-capacity models maintain stable system rankings across personas, whereas smaller models significantly amplify instability. This sensitivity is particularly pronounced in neural rerankers and RAG systems, with persona role proving more influential than persona source, thereby offering a novel diagnostic tool for assessing the robustness of LLM-based evaluation.
📝 Abstract
Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about how assessor framing affects judgment reliability and downstream system comparison. We study persona conditioning as a diagnostic mechanism for exposing LLM assessor sensitivity. Using task-oriented personas drawn from two complementary sources (PersonaHub and NVIDIA Nemotron-Personas-USA), we instantiate five assessor roles emphasizing intent interpretation, domain expertise, contrastive judgment, evidence verification, and global search-quality assessment, compared with a standard UMBRELA baseline. Across six LLM backbones on TREC DL20 and RAG24, our analyses reveal structured rather than uniform assessor sensitivity. Judgments usually remain close to the baseline while shifting assessment strictness, evidential threshold, or interpretation emphasis rather than producing widespread relevance reversals. At the system level, high-capacity models preserve system-ranking agreement, while smaller models amplify persona-induced instability. Local rank-displacement analysis shows sensitivity concentrates on particular retrieval systems and system types, especially neural ranking/reranking systems on DL20 and RAG-oriented pipelines on RAG24. Persona source matters less than assessor role and model capacity. These findings position persona-conditioned judging as a controlled sensitivity probe for stress-testing LLM-based IR evaluation pipelines and identifying systems whose evaluation outcomes are sensitive to assessor framing.