A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset
This study addresses the lack of a realistic evaluation benchmark for Greek-language book retrieval by introducing CUP, the first dataset comprising 868 Greek bibliographic records and 104 expert-annotated queries. The authors systematically evaluate sparse (BM25), dense (sentence-transformers), hybrid, and large language model (LLM)-augmented retrieval approaches. Experimental results demonstrate that hybrid retrieval achieves the best overall performance; BM25 excels on named entity queries, while dense and hybrid methods substantially improve effectiveness on natural language, noisy, cross-lingual, and conceptual queries. Multilingual embeddings consistently outperform monolingual models, and LLM-based post-processing yields gains at a higher computational cost. This work establishes the first fine-grained retrieval benchmark for Greek publishing and provides a comprehensive comparison of modern retrieval strategies in this underexplored domain.