🤖 AI Summary
This study addresses the limitations of holistic scoring in metaphor explanation evaluation, which often overlooks quality structure and human disagreement. We propose a cognition-driven, six-dimensional assessment framework to capture these nuances. Through large-scale annotation and clustering analysis, we reveal the multidimensionality of explanation quality and systematic patterns of disagreement, validating that an automated evaluation pipeline can effectively recover this structure. Our results demonstrate that automatic models can predict key dimensions, with prediction errors significantly correlating with human disagreement. By overcoming the constraints of single-score metrics, this work establishes a fine-grained, diagnostically valuable evaluation paradigm for open-ended generation tasks, offering deeper insights into model performance and human alignment.
📝 Abstract
Current evaluation of metaphor explanations relies mainly on holistic quality ratings, revealing little about how explanation quality is structured or where human judgments agree and diverge. We introduce a cognitively motivated framework that decomposes metaphor explanation quality into six theoretically grounded dimensions. In a dense annotation study (11,200 ratings), we find that: {\bfseries(i)} explanation quality is genuinely multidimensional; {\bfseries(ii)} annotator disagreement is systematic rather than random; and {\bfseries(iii)} the six dimensions collapse into a shared cluster and two independent axes of judgment. An exploratory feasibility study further shows that a standard automatic evaluation pipeline can recover parts of this structure, predicting the most discriminative dimensions well while its errors correlate human (dis)agreement. Together, these results suggest that multidimensional evaluation offers richer diagnostic insight than holistic ratings, and that automatic evaluators for open-ended generation tasks should be judged on how well they preserve the structure of human judgment.