SEval-Ex: A Statement-Level Framework for Explainable Summarization Evaluation

📅 2025-05-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Addressing the trade-off between performance and interpretability in summary evaluation, this paper proposes a fine-grained assessment framework based on atomic sentence decomposition: leveraging large language models (LLMs) to perform sentence-level extraction and semantic alignment between source documents and summaries, thereby constructing traceable and verifiable decision evidence chains. The method introduces the first sentence-level alignment evaluation paradigm, overcoming the limitations of conventional summary-level scoring. It is the first approach to achieve state-of-the-art (SOTA) performance while providing transparent, step-by-step reasoning paths and demonstrating robustness against hallucination. On the SummEval benchmark, it attains a consistency score of 0.580—significantly outperforming the GPT-4 evaluator (0.521)—thereby advancing both evaluation accuracy and interpretability.

Technology Category

Application Category

📝 Abstract
Evaluating text summarization quality remains a critical challenge in Natural Language Processing. Current approaches face a trade-off between performance and interpretability. We present SEval-Ex, a framework that bridges this gap by decomposing summarization evaluation into atomic statements, enabling both high performance and explainability. SEval-Ex employs a two-stage pipeline: first extracting atomic statements from text source and summary using LLM, then a matching between generated statements. Unlike existing approaches that provide only summary-level scores, our method generates detailed evidence for its decisions through statement-level alignments. Experiments on the SummEval benchmark demonstrate that SEval-Ex achieves state-of-the-art performance with 0.580 correlation on consistency with human consistency judgments, surpassing GPT-4 based evaluators (0.521) while maintaining interpretability. Finally, our framework shows robustness against hallucination.
Problem

Research questions and friction points this paper is trying to address.

Bridging performance-interpretability gap in summarization evaluation
Providing statement-level evidence for evaluation decisions
Achieving robust evaluation against hallucination in summaries
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decomposes evaluation into atomic statements for explainability
Uses LLM for statement extraction and matching
Achieves state-of-the-art performance with interpretability
💼 Related Jobs
No related jobs found.
T
Tanguy Herserant
AgroParisTech - MIA, 22 place de l'Agronomie, 91120 Palaiseau, France
V
Vincent Guigue
AgroParisTech - MIA, 22 place de l'Agronomie, 91120 Palaiseau, France