LLM-based relevance assessment still can't replace human relevance assessment
This work challenges the feasibility of replacing human assessors with large language models (LLMs) for relevance evaluation in information retrieval. Method: Leveraging empirical analysis on TREC 2024 data, adversarial system submissions, theoretical modeling, and robustness testing, the study systematically investigates LLM-based relevance assessment. Contribution/Results: It identifies, for the first time, an “intrinsic narcissism” in LLM evaluation—where assessments rely on self-referential generative logic, inducing susceptibility to metric manipulation, self-referential bias, and overfitting. Experiments demonstrate that targeted optimization can artificially inflate LLM scores, and their judgments fail to support sustainable iterative improvement of retrieval systems. The findings establish fundamental deficiencies in both theoretical reliability and practical robustness of LLM-based evaluation, reaffirming the irreplaceable role of human assessment. This work provides a critical caution and methodological reflection for retrieval evaluation paradigms.