Institution profile

Elsevier

Industry researcheurope · nl
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

DeepImagine: Learning Biomedical Reasoning via Successive Counterfactual Imagining

Apr 24, 2026

Current large language models struggle to capture the underlying causal mechanisms when predicting clinical trial outcomes. To address this limitation, this work proposes a novel training paradigm that integrates supervised fine-tuning, reinforcement learning guided by validation-based rewards, and synthetic counterfactual reasoning trajectories. By leveraging counterfactual pairs to construct training data, the approach steers models—particularly those under 10B parameters, such as Qwen3.5-9B—toward learning interpretable biomedical causal reasoning processes through iterative counterfactual imagination. The method substantially outperforms both unadapted language models and conventional correlation-based baselines, while simultaneously generating transparent and human-interpretable reasoning pathways that reflect plausible causal structures in clinical contexts.

0 citationsRead paper

Detecting Data Contamination in Large Language Models

Apr 21, 2026

This study addresses the challenge of membership inference for copyrighted or sensitive content in the training data of large language models (LLMs) under black-box settings. To overcome the lack of standardized evaluation in existing approaches, the authors propose a novel "familiarity ranking" method that enhances model output flexibility to better reveal memorization tendencies toward specific data points. The work systematically evaluates multiple black-box membership inference attacks on mainstream LLMs within a unified benchmark. Experimental results demonstrate that all tested methods struggle to reliably infer membership, achieving near-random performance (AUC-ROC ≈ 0.5). Moreover, the stronger generalization capabilities of advanced models further exacerbate the difficulty of black-box detection, highlighting fundamental limitations of current techniques for this task.

0 citationsRead paper

CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical Trial Outcome Prediction

Apr 17, 2026

This work addresses the longstanding lack of reliable, uncontaminated benchmarks for evaluating clinical trial outcome prediction, which has hindered rigorous assessment of AI systems’ ability to forecast real-world future events. To this end, we introduce CT Open—an open, real-time evaluation platform that hosts four annual challenges and enforces strict timestamping to ensure predictions are submitted before outcomes become public. We develop the first fully automated decontamination pipeline, combining large language model–driven iterative web search with expert annotation to accurately determine the earliest public disclosure time of trial results. The project releases a training dataset alongside two temporally anchored test benchmarks—Winter 2025 and Summer 2025—establishing a fair, reproducible framework for evaluating AI-driven forecasting of real-world clinical events.

0 citationsRead paper

AnalyticsGPT: An LLM Workflow for Scientometric Question Answering

Feb 10, 2026

This work proposes the first end-to-end system based on large language models (LLMs) to address the challenges of automated question answering for metascientific inquiries in scientometrics, such as academic entity recognition and multidimensional metric retrieval. By integrating retrieval-augmented generation (RAG) with an agent-based architecture, the system autonomously decomposes complex questions, plans analytical workflows, retrieves relevant data, and generates structured, high-level analyses. The approach innovatively leverages the reasoning and task-planning capabilities of LLMs within the domain of scientometrics and incorporates a proprietary research performance database. Expert evaluation and LLM-as-judge validation demonstrate that the system efficiently produces accurate, logically coherent analytical reports suitable for advanced scientometric inquiry.

0 citationsRead paper

LiRA: A Multi-Agent Framework for Reliable and Readable Literature Review Generation

Oct 01, 2025

The exponential growth of scientific literature impedes systematic reviews from simultaneously achieving comprehensiveness, timeliness, and readability; existing automated approaches predominantly focus on retrieval and screening, neglecting factual accuracy and textual quality during the writing phase. This paper introduces the first multi-agent collaborative framework specifically designed for review authoring, which emulates human workflow through four specialized agents—Outline, Writing, Editing, and Reviewing—that jointly perform end-to-end generation. The framework requires no domain-specific fine-tuning and integrates systematic literature retrieval, structured citation handling, and segment-wise controllable text generation. Evaluated on SciReviewGen and ScienceDirect benchmarks, it significantly outperforms baselines including AutoSurvey and MASS-Survey. Generated reviews achieve near-human performance in factual accuracy, linguistic readability, and citation conformity.

0 citationsRead paper
Recent publications

Latest Papers

DeepImagine: Learning Biomedical Reasoning via Successive Counterfactual Imagining

Apr 24, 2026

Current large language models struggle to capture the underlying causal mechanisms when predicting clinical trial outcomes. To address this limitation, this work proposes a novel training paradigm that integrates supervised fine-tuning, reinforcement learning guided by validation-based rewards, and synthetic counterfactual reasoning trajectories. By leveraging counterfactual pairs to construct training data, the approach steers models—particularly those under 10B parameters, such as Qwen3.5-9B—toward learning interpretable biomedical causal reasoning processes through iterative counterfactual imagination. The method substantially outperforms both unadapted language models and conventional correlation-based baselines, while simultaneously generating transparent and human-interpretable reasoning pathways that reflect plausible causal structures in clinical contexts.

0 citationsRead paper

Detecting Data Contamination in Large Language Models

Apr 21, 2026

This study addresses the challenge of membership inference for copyrighted or sensitive content in the training data of large language models (LLMs) under black-box settings. To overcome the lack of standardized evaluation in existing approaches, the authors propose a novel "familiarity ranking" method that enhances model output flexibility to better reveal memorization tendencies toward specific data points. The work systematically evaluates multiple black-box membership inference attacks on mainstream LLMs within a unified benchmark. Experimental results demonstrate that all tested methods struggle to reliably infer membership, achieving near-random performance (AUC-ROC ≈ 0.5). Moreover, the stronger generalization capabilities of advanced models further exacerbate the difficulty of black-box detection, highlighting fundamental limitations of current techniques for this task.

0 citationsRead paper

CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical Trial Outcome Prediction

Apr 17, 2026

This work addresses the longstanding lack of reliable, uncontaminated benchmarks for evaluating clinical trial outcome prediction, which has hindered rigorous assessment of AI systems’ ability to forecast real-world future events. To this end, we introduce CT Open—an open, real-time evaluation platform that hosts four annual challenges and enforces strict timestamping to ensure predictions are submitted before outcomes become public. We develop the first fully automated decontamination pipeline, combining large language model–driven iterative web search with expert annotation to accurately determine the earliest public disclosure time of trial results. The project releases a training dataset alongside two temporally anchored test benchmarks—Winter 2025 and Summer 2025—establishing a fair, reproducible framework for evaluating AI-driven forecasting of real-world clinical events.

0 citationsRead paper

AnalyticsGPT: An LLM Workflow for Scientometric Question Answering

Feb 10, 2026

This work proposes the first end-to-end system based on large language models (LLMs) to address the challenges of automated question answering for metascientific inquiries in scientometrics, such as academic entity recognition and multidimensional metric retrieval. By integrating retrieval-augmented generation (RAG) with an agent-based architecture, the system autonomously decomposes complex questions, plans analytical workflows, retrieves relevant data, and generates structured, high-level analyses. The approach innovatively leverages the reasoning and task-planning capabilities of LLMs within the domain of scientometrics and incorporates a proprietary research performance database. Expert evaluation and LLM-as-judge validation demonstrate that the system efficiently produces accurate, logically coherent analytical reports suitable for advanced scientometric inquiry.

0 citationsRead paper

LiRA: A Multi-Agent Framework for Reliable and Readable Literature Review Generation

Oct 01, 2025

The exponential growth of scientific literature impedes systematic reviews from simultaneously achieving comprehensiveness, timeliness, and readability; existing automated approaches predominantly focus on retrieval and screening, neglecting factual accuracy and textual quality during the writing phase. This paper introduces the first multi-agent collaborative framework specifically designed for review authoring, which emulates human workflow through four specialized agents—Outline, Writing, Editing, and Reviewing—that jointly perform end-to-end generation. The framework requires no domain-specific fine-tuning and integrates systematic literature retrieval, structured citation handling, and segment-wise controllable text generation. Evaluated on SciReviewGen and ScienceDirect benchmarks, it significantly outperforms baselines including AutoSurvey and MASS-Survey. Generated reviews achieve near-human performance in factual accuracy, linguistic readability, and citation conformity.

0 citationsRead paper