π€ AI Summary
This study addresses the limitation of existing summarization evaluation metrics in overlooking individual needs and failing to measure utility for specific readers. We propose a novel reader-centric dimension termed "information satisfaction," integrating reader profiles as stable signals to construct a personalized evaluation framework. This framework is validated through perturbation testing, expert assessment, and LLM-as-judge methodologies. Our findings demonstrate that both traditional metrics and mainstream LLM-based evaluators exhibit low correlation with information satisfaction and poor alignment with human judgment, revealing significant limitations in personalized assessment. By bridging the gap in measuring user-specific information fulfillment, this work establishes a new paradigm for summarization evaluation that prioritizes individual reader utility over generic quality standards.
π Abstract
The majority of work on summarization evaluation focuses on general summary quality (e.g., ROUGE, BERTScore) or specific desired properties (e.g., readability, factuality). However, these metrics fail to measure the utility of a summary to an individual user. For example, a biomedical researcher learning about the latest vaccine research will have different informational needs from a family doctor. Query-focused summarization captures part of this need, but in practice, users rarely state everything relevant in a query: a single short query is likely inadequate to distinguish the needs of a researcher from those of a physician. By contrast, a reader's background or persona (their role and expertise) is comparatively stable across queries and recovers much of this missing context, which makes it a practical signal for assessing whether a summary satisfies that reader's needs. In this work, we assess how sensitive popular summarization metrics are to both informational and persona differences, and find that many popular metrics, including strong LLM-as-judge metrics, fail basic perturbation tests of informational content. We additionally conduct an expert human evaluation, measuring summary preferences based on information satisfaction given a specific person's background and use case. We find that both traditional and LLM-based metrics are insufficient measures of information satisfaction and agree poorly with human judgment.