BookAsSumQA: An Evaluation Framework for Aspect-Based Book Summarization via Question Answering
Aspect-based summarization of long texts—such as books—suffers from a scarcity of high-quality reference summaries, prohibitively high costs of human evaluation, and poor scalability. Method: This paper proposes BookAsSumQA, the first automated evaluation framework for aspect-based summarization of long literary texts. It leverages narrative knowledge graphs to automatically generate aspect-specific question-answer (QA) pairs, eliminating the need for manually annotated reference summaries. By integrating large language models (LLMs) with retrieval-augmented generation (RAG), it uses QA accuracy as a proxy metric for summary quality. Contribution/Results: Experiments demonstrate that BookAsSumQA effectively discriminates among diverse summarization methods. Notably, it is the first to empirically reveal RAG’s significant superiority over LLM-only approaches in long-document aspect-based summarization. The framework is both scalable and practically applicable, enabling efficient, reference-free evaluation.