LLM-based relevance assessment still can't replace human relevance assessment

📅 2024-12-22
🏛️ arXiv.org
📈 Citations: 11
Influential: 1
📄 PDF
🤖 AI Summary
This work challenges the feasibility of replacing human assessors with large language models (LLMs) for relevance evaluation in information retrieval. Method: Leveraging empirical analysis on TREC 2024 data, adversarial system submissions, theoretical modeling, and robustness testing, the study systematically investigates LLM-based relevance assessment. Contribution/Results: It identifies, for the first time, an “intrinsic narcissism” in LLM evaluation—where assessments rely on self-referential generative logic, inducing susceptibility to metric manipulation, self-referential bias, and overfitting. Experiments demonstrate that targeted optimization can artificially inflate LLM scores, and their judgments fail to support sustainable iterative improvement of retrieval systems. The findings establish fundamental deficiencies in both theoretical reliability and practical robustness of LLM-based evaluation, reaffirming the irreplaceable role of human assessment. This work provides a critical caution and methodological reflection for retrieval evaluation paradigms.

Technology Category

Application Category

📝 Abstract
The use of large language models (LLMs) for relevance assessment in information retrieval has gained significant attention, with recent studies suggesting that LLM-based judgments provide comparable evaluations to human judgments. Notably, based on TREC 2024 data, Upadhyay et al. make a bold claim that LLM-based relevance assessments, such as those generated by the UMBRELA system, can fully replace traditional human relevance assessments in TREC-style evaluations. This paper critically examines this claim, highlighting practical and theoretical limitations that undermine the validity of this conclusion. First, we question whether the evidence provided by Upadhyay et al. really supports their claim, particularly if a test collection is used asa benchmark for future improvements. Second, through a submission deliberately intended to do so, we demonstrate the ease with which automatic evaluation metrics can be subverted, showing that systems designed to exploit these evaluations can achieve artificially high scores. Theoretical challenges -- such as the inherent narcissism of LLMs, the risk of overfitting to LLM-based metrics, and the potential degradation of future LLM performance -- must be addressed before LLM-based relevance assessments can be considered a viable replacement for human judgments.
Problem

Research questions and friction points this paper is trying to address.

Examines if LLM-based relevance assessments can replace human judgments
Highlights risks of subverting automatic evaluation metrics artificially
Addresses theoretical challenges like LLM narcissism and overfitting risks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Questioning LLM-based relevance assessment validity
Demonstrating automatic evaluation metrics vulnerability
Addressing theoretical LLM limitations