Institution profile

Institute of Computer Science

Academic institutioneurope · gr
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset

Jul 23, 2026

This study addresses the lack of a realistic evaluation benchmark for Greek-language book retrieval by introducing CUP, the first dataset comprising 868 Greek bibliographic records and 104 expert-annotated queries. The authors systematically evaluate sparse (BM25), dense (sentence-transformers), hybrid, and large language model (LLM)-augmented retrieval approaches. Experimental results demonstrate that hybrid retrieval achieves the best overall performance; BM25 excels on named entity queries, while dense and hybrid methods substantially improve effectiveness on natural language, noisy, cross-lingual, and conceptual queries. Multilingual embeddings consistently outperform monolingual models, and LLM-based post-processing yields gains at a higher computational cost. This work establishes the first fine-grained retrieval benchmark for Greek publishing and provides a comprehensive comparison of modern retrieval strategies in this underexplored domain.

0 citationsRead paper

Exploring Gender Bias Beyond Occupational Titles

Jul 03, 2025

This study uncovers implicit gender bias in language that extends beyond occupational stereotypes, focusing on non-occupational linguistic elements—particularly action verbs and object nouns. Method: We propose the first fine-grained, multilingual (including Japanese) gender bias evaluation framework, comprising the GenderLexicon dataset and an interpretable, context-aware bias quantification model. Unlike conventional occupation-based bias detection, our approach systematically identifies and attributes gender bias to dynamic semantic units such as verb–noun collocations. Contribution/Results: Experiments across five cross-lingual, multi-domain datasets demonstrate that such bias is pervasive in everyday syntactic and semantic structures. Our model achieves strong generalizability and interpretability, enabling precise bias localization and causal attribution. This work establishes a novel paradigm for bias溯源 (traceability) and mitigation, advancing both computational linguistics and fairness-aware NLP.

0 citationsRead paper
Recent publications

Latest Papers

A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset

Jul 23, 2026

This study addresses the lack of a realistic evaluation benchmark for Greek-language book retrieval by introducing CUP, the first dataset comprising 868 Greek bibliographic records and 104 expert-annotated queries. The authors systematically evaluate sparse (BM25), dense (sentence-transformers), hybrid, and large language model (LLM)-augmented retrieval approaches. Experimental results demonstrate that hybrid retrieval achieves the best overall performance; BM25 excels on named entity queries, while dense and hybrid methods substantially improve effectiveness on natural language, noisy, cross-lingual, and conceptual queries. Multilingual embeddings consistently outperform monolingual models, and LLM-based post-processing yields gains at a higher computational cost. This work establishes the first fine-grained retrieval benchmark for Greek publishing and provides a comprehensive comparison of modern retrieval strategies in this underexplored domain.

0 citationsRead paper

Exploring Gender Bias Beyond Occupational Titles

Jul 03, 2025

This study uncovers implicit gender bias in language that extends beyond occupational stereotypes, focusing on non-occupational linguistic elements—particularly action verbs and object nouns. Method: We propose the first fine-grained, multilingual (including Japanese) gender bias evaluation framework, comprising the GenderLexicon dataset and an interpretable, context-aware bias quantification model. Unlike conventional occupation-based bias detection, our approach systematically identifies and attributes gender bias to dynamic semantic units such as verb–noun collocations. Contribution/Results: Experiments across five cross-lingual, multi-domain datasets demonstrate that such bias is pervasive in everyday syntactic and semantic structures. Our model achieves strong generalizability and interpretability, enabling precise bias localization and causal attribution. This work establishes a novel paradigm for bias溯源 (traceability) and mitigation, advancing both computational linguistics and fairness-aware NLP.

0 citationsRead paper