GhazalBench: Usage-Grounded Evaluation of LLMs on Persian Ghazals
While current large language models demonstrate a capacity to comprehend the poetic essence of Persian ghazal poetry, they struggle to reproduce its culturally normative surface form in open-ended generation. This work proposes GhazalBench, a novel evaluation benchmark that, for the first time, incorporates the ability to generate culturally conformant textual forms as a core assessment dimension. The benchmark introduces two tasks: prose-to-poetry comprehension and cue-guided reconstruction of normative verses, complemented by a comparative experiment using English sonnets. Findings reveal that mainstream multilingual models generally excel at semantic understanding but underperform in open-generation of structurally and culturally compliant ghazals. Discriminative tasks notably narrow the performance gap, and the observed limitations are primarily attributed to insufficient coverage of relevant training data rather than inherent architectural constraints.