Institution profile

Athena RC

Academic institutioneurope · gr
Official website
Research library17linked papers
Opportunities0open roles
Selected work

Representative Papers

Scientific Knowledge Discovery in the Age of Large Language Models

Jul 29, 2026

This study addresses the inefficiency of traditional literature retrieval methods, which rely heavily on manually crafted queries and screening, in the face of exponentially growing academic publications. The authors present a systematic review of 34 peer-reviewed studies that apply generative large language models (LLMs) to scientific literature retrieval and screening, offering the first comprehensive mapping of their application paradigms in scientific knowledge discovery. By leveraging Boolean queries on the OpenAIRE Graph and analyzing approaches through the lenses of prompt engineering, model adaptation, and architectural design, the work identifies key technical pathways and evaluation frameworks. Beyond structuring the current landscape, the study highlights the substantial potential of LLMs to enhance the automation and efficiency of scientific discovery, providing a systematic reference for future research.

0 citationsRead paper

ANN Search: Recall What Matters

Jun 03, 2026

This work addresses the limitations of Recall@k as the dominant evaluation metric in approximate nearest neighbor (ANN) search, which often overestimates retrieval quality and incurs redundant computation. The authors propose 1/Ratio@k—an inverse approximation ratio that is hyperparameter-free and directly computable from benchmark data—as a more principled alternative. Through extensive evaluation of state-of-the-art ANN algorithms across diverse high-dimensional datasets, coupled with efficiency analysis and validation on downstream tasks such as classification and retrieval-augmented generation, the study demonstrates that optimizing 1/Ratio@k significantly reduces computational overhead while preserving practical utility. Moreover, 1/Ratio@k exhibits substantially stronger correlation with real-world effectiveness—measured by label accuracy and semantic similarity—than Recall@k.

0 citationsRead paper

Enhancing Scientific Discourse: Machine Translation for the Scientific Domain

May 20, 2026

This study addresses the linguistic barriers impeding the global dissemination of scientific research, which generic machine translation systems struggle to overcome due to their inability to accurately handle domain-specific terminology and complex syntactic structures in scholarly texts. To bridge this gap, the authors present the first systematic construction of Spanish–English, French–English, and Portuguese–English parallel and monolingual corpora spanning four scientific subfields: cancer, energy, neuroscience, and transportation. Leveraging these resources, they perform domain-adaptive fine-tuning of neural machine translation models. Experimental results demonstrate that the fine-tuned systems significantly outperform generic baselines in translation quality for scientific content, thereby confirming the critical role of multilingual, multidisciplinary specialized corpora in enhancing the accuracy and fluency of research literature translation.

0 citationsRead paper
Recent publications

Latest Papers

Scientific Knowledge Discovery in the Age of Large Language Models

Jul 29, 2026

This study addresses the inefficiency of traditional literature retrieval methods, which rely heavily on manually crafted queries and screening, in the face of exponentially growing academic publications. The authors present a systematic review of 34 peer-reviewed studies that apply generative large language models (LLMs) to scientific literature retrieval and screening, offering the first comprehensive mapping of their application paradigms in scientific knowledge discovery. By leveraging Boolean queries on the OpenAIRE Graph and analyzing approaches through the lenses of prompt engineering, model adaptation, and architectural design, the work identifies key technical pathways and evaluation frameworks. Beyond structuring the current landscape, the study highlights the substantial potential of LLMs to enhance the automation and efficiency of scientific discovery, providing a systematic reference for future research.

0 citationsRead paper

ANN Search: Recall What Matters

Jun 03, 2026

This work addresses the limitations of Recall@k as the dominant evaluation metric in approximate nearest neighbor (ANN) search, which often overestimates retrieval quality and incurs redundant computation. The authors propose 1/Ratio@k—an inverse approximation ratio that is hyperparameter-free and directly computable from benchmark data—as a more principled alternative. Through extensive evaluation of state-of-the-art ANN algorithms across diverse high-dimensional datasets, coupled with efficiency analysis and validation on downstream tasks such as classification and retrieval-augmented generation, the study demonstrates that optimizing 1/Ratio@k significantly reduces computational overhead while preserving practical utility. Moreover, 1/Ratio@k exhibits substantially stronger correlation with real-world effectiveness—measured by label accuracy and semantic similarity—than Recall@k.

0 citationsRead paper

Enhancing Scientific Discourse: Machine Translation for the Scientific Domain

May 20, 2026

This study addresses the linguistic barriers impeding the global dissemination of scientific research, which generic machine translation systems struggle to overcome due to their inability to accurately handle domain-specific terminology and complex syntactic structures in scholarly texts. To bridge this gap, the authors present the first systematic construction of Spanish–English, French–English, and Portuguese–English parallel and monolingual corpora spanning four scientific subfields: cancer, energy, neuroscience, and transportation. Leveraging these resources, they perform domain-adaptive fine-tuning of neural machine translation models. Experimental results demonstrate that the fine-tuned systems significantly outperform generic baselines in translation quality for scientific content, thereby confirming the critical role of multilingual, multidisciplinary specialized corpora in enhancing the accuracy and fluency of research literature translation.

0 citationsRead paper