NormasTCU --- A Brazilian Portuguese IR Dataset and an Evaluation of LLM-as-a-Judge for Relevance Assessment

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决葡萄牙语信息检索缺乏公共数据集及专业领域相关性评估成本高的问题,通过构建NormasTCU数据集并使用大语言模型作为评判工具进行相关性评估。
📝 Abstract
Portuguese Information Retrieval (IR) lacks public datasets, and relevance assessment for specialized collections remains costly. While Large Language Models (LLMs) increasingly support relevance assessment, their reliability in non-English specialized domains remains unclear. We introduce NormasTCU (https://huggingface.co/datasets/LeandroRibeiro/NormasTCU), a Brazilian Portuguese IR dataset with 14,469 legal documents, 46 queries, and 3,048 human judgments over 812 query-document pairs. Using NormasTCU, we evaluated LLM-as-a-judge for relevance assessment by prompting three models with two prompt techniques to grade these pairs. We then compared the rankings of 15 IR systems derived from LLM-generated and human reference qrels. LLMs consistently showed a positive scoring bias (mean absolute error: 0.46--0.66 on a 0-2 scale). Furthermore, pair-level agreement with human judgments achieved only fair to moderate levels, with Cohen's kappa ranging from 0.32 to 0.53. Despite this bias, LLM-generated judgments often yielded highly similar system rankings for nDCG@10 and MRR (observed Kendall's tau greater than or equal 0.90, although the bootstrap confidence intervals did not always remain above this threshold), but were less reliable for P@10 and R@10. Notably, LLM-based rankings were sometimes more strongly correlated with the reference ranking than individual human annotations were. As a practical implication, our results suggest that LLMs could effectively support scalable relevance assessment in specialized Portuguese corpora when evaluated using nDCG or MRR (rank-aware metrics), but they should be avoided when relying on precision or recall.
Problem

Research questions and friction points this paper is trying to address.

Portuguese Information Retrieval
relevance assessment
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-as-a-judge
relevance assessment
Brazilian Portuguese IR dataset
nDCG@10 and MRR
💼 Related Jobs
No related jobs found.
Leandro Carísio Fernandes
Leandro Carísio Fernandes
Câmara dos Deputados
M
Marcus Vinícius Borela de Castro
Tribunal de Contas da União (TCU), Brasília, Brazil.
L
Leandro dos Santos Ribeiro
Tribunal de Contas da União (TCU), Brasília, Brazil.
L
Leonardo Augusto da Silva Pacheco
Tribunal de Contas da União (TCU), Brasília, Brazil.
E
Edans Flávius de Oliveira Sandes
Tribunal de Contas da União (TCU), Brasília, Brazil.