π€ AI Summary
This study addresses the longstanding reliance on subjective judgment in narrative quality assessment by introducing a computational framework grounded in 33 quantifiable linguistic features spanning lexical, syntactic, and semantic dimensions. For the first time, this work systematically applies multidimensional quantitative stylometric indicators to the automatic evaluation of narrative quality. Leveraging natural language processing, clustering analysis, and similarity matrix construction, the proposed model achieves near-perfect discrimination between texts authored by professional editors and self-published writers. Furthermore, it significantly outperforms existing evaluation metrics on a manually annotated dataset, thereby overcoming the limitations inherent in traditional story-level assessment approaches.
π Abstract
The evaluation of narrative quality remains a complex challenge, as it involves subjective factors such as plot, character development, and emotional impact. This work proposes a quantitative approach to narrative assessment by focusing on the linguistic dimension as a primary indicator of quality. The paper presents a methodology for the automatic evaluation of narrative based on the extraction of a comprehensive set of 33 quantitative linguistic features categorized into lexical, syntactic, and semantic groups. To test the model, an experiment was conducted on a specialized corpus of 23 books, including canonical masterpieces and self-published works. Through a similarity matrix, the system successfully clustered the narratives, distinguishing almost perfectly between professionally edited and self-published texts. Furthermore, the methodology was validated against a human-annotated dataset; it significantly outperforms traditional story-level evaluation metrics, demonstrating the effectiveness of quantitative linguistic features in assessing narrative quality.