FactSim: Fact-Checking for Opinion Summarization
This work addresses the challenge that existing automatic evaluation metrics struggle to accurately assess the factual consistency of opinion summaries generated by large language models. To this end, the authors propose FactSim, an end-to-end fully automated evaluation method that extracts factual claims from both the generated summary and the original user reviews, and introduces a robust fact similarity scoring mechanism designed to handle negations, paraphrases, and elaborations. This approach effectively measures both factual consistency and coverage. Experimental results demonstrate that FactSim achieves significantly higher correlation with human judgments than current state-of-the-art automatic metrics, offering a more reliable proxy for human evaluation in assessing factual faithfulness.