Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文提出预测驱动评估框架,结合少量人工判断和大规模自动评分,解决非验证任务中人工评估成本高、自动度量有偏的问题。
📝 Abstract
Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction-powered inference (PPI), we propose prediction-powered evaluation, a framework that combines limited human judgments with large-scale automatic scores to obtain data-efficient system comparisons that are provably unbiased. We develop parametric and non-parametric procedures, analyze the efficiency trade-off between paired and unpaired designs, and validate the framework on six WMT datasets. We further introduce the Prediction-Powered Saving Ratio (PPSR), a meta-metric that measures how much human annotation an automatic metric can save when used within prediction-powered evaluation. PPSR directly targets metric utility for prediction-powered evaluation and yields more discriminative and stable metric rankings than existing system-level meta-metrics. Overall, our new paradigm reframes automatic metrics as tools for reducing human annotation cost rather than replacing human judgment, and applies broadly to non-verifiable tasks.
Problem

Research questions and friction points this paper is trying to address.

human evaluation
automatic metrics
unbiased comparison
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prediction-Powered Evaluation
Prediction-Powered Saving Ratio (PPSR)
Unbiased System Comparisons