Institution profile

Bayer AG

Industry researcheurope · de
Official website
Research library24linked papers
Opportunities0open roles
Selected work

Representative Papers

Agents in the Wild: Where Research Meets Deployment

Jul 21, 2026

This work addresses critical challenges in the real-world deployment of large language model (LLM)-driven agent systems, particularly concerning robustness, safety, and reliability. Bridging academic advances with industrial practice, the study presents case studies from software engineering, scientific discovery, and finance to distill reusable design patterns and an evaluation checklist. It integrates key techniques including LLM-based reasoning and planning, multi-agent coordination, validation pipelines, fallback mechanisms, and human-in-the-loop oversight. The proposed cross-domain deployment framework has been validated in pharmaceutical discovery and financial systems, demonstrating significant improvements in stability and trustworthiness of agent systems in real-world settings, thereby narrowing the gap between research innovation and practical implementation.

0 citationsRead paper

Episodic-Semantic Memory Architecture for Long-Horizon Scientific Agents

May 17, 2026

This work addresses the cognitive degradation and escalating computational overhead in large language models during extended scientific collaboration, caused by context saturation. To overcome these limitations, we propose a dual-process memory architecture that decouples short-term episodic memory (fixed to the latest 10 messages) from long-term semantic knowledge (growing at approximately 3 tokens per message). The framework integrates domain-specific knowledge compression, dual-channel episodic-semantic memory, and cross-model verification, enabling robust handling of parameter contradictions, multi-hop reasoning across collaboration stages, and precise retention of technical facts. Evaluated across six mainstream large language models, our system maintains 70–85% accuracy over 15,000 messages with only 1–2 seconds of latency, reduces token consumption by 62%, and successfully manages over 14,000 scientific facts (125k tokens), substantially surpassing the capacity and efficiency limits of conventional full-context approaches.

0 citationsRead paper

Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets

May 14, 2026

This work addresses the limitations of existing Croissant metadata generation approaches, which rely on public platforms and struggle to accommodate governed or large-scale local datasets. The authors propose the first open-source, local-first command-line tool that directly generates Croissant-compliant JSON-LD metadata from local directories via a modular processor registration mechanism, supporting mainstream formats such as Parquet. By eliminating dependence on external platforms, this method significantly enhances the discoverability and reusability of private, high-value datasets. Experimental evaluation across more than 140 datasets—including MIMIC-IV with 886 million rows—demonstrates that the generated metadata achieves 97–100% accuracy, matching or exceeding that of manual curation or standard methods.

0 citationsRead paper

Communicating results in trials with multiple hypotheses or adaptive design features

May 05, 2026

This study addresses the underappreciated challenge of estimation and communication following multiplicity adjustment within the frequentist framework in complex clinical trials, where multiple endpoints, interim data looks, or group comparisons often introduce estimation bias and complicate interpretation, thereby undermining transparency in benefit–risk assessment. By integrating advanced methodologies such as adaptive designs and graphical approaches to multiple testing, the work illustrates through concrete examples the limitations of current strategies in conveying trial results meaningfully. The research underscores the need to critically reevaluate prevailing practices and foster interdisciplinary dialogue to enhance both the accuracy of effect estimation and the clarity of result communication, ultimately informing future methodological standards and regulatory guidance.

0 citationsRead paper
Recent publications

Latest Papers

Agents in the Wild: Where Research Meets Deployment

Jul 21, 2026

This work addresses critical challenges in the real-world deployment of large language model (LLM)-driven agent systems, particularly concerning robustness, safety, and reliability. Bridging academic advances with industrial practice, the study presents case studies from software engineering, scientific discovery, and finance to distill reusable design patterns and an evaluation checklist. It integrates key techniques including LLM-based reasoning and planning, multi-agent coordination, validation pipelines, fallback mechanisms, and human-in-the-loop oversight. The proposed cross-domain deployment framework has been validated in pharmaceutical discovery and financial systems, demonstrating significant improvements in stability and trustworthiness of agent systems in real-world settings, thereby narrowing the gap between research innovation and practical implementation.

0 citationsRead paper

Episodic-Semantic Memory Architecture for Long-Horizon Scientific Agents

May 17, 2026

This work addresses the cognitive degradation and escalating computational overhead in large language models during extended scientific collaboration, caused by context saturation. To overcome these limitations, we propose a dual-process memory architecture that decouples short-term episodic memory (fixed to the latest 10 messages) from long-term semantic knowledge (growing at approximately 3 tokens per message). The framework integrates domain-specific knowledge compression, dual-channel episodic-semantic memory, and cross-model verification, enabling robust handling of parameter contradictions, multi-hop reasoning across collaboration stages, and precise retention of technical facts. Evaluated across six mainstream large language models, our system maintains 70–85% accuracy over 15,000 messages with only 1–2 seconds of latency, reduces token consumption by 62%, and successfully manages over 14,000 scientific facts (125k tokens), substantially surpassing the capacity and efficiency limits of conventional full-context approaches.

0 citationsRead paper

Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets

May 14, 2026

This work addresses the limitations of existing Croissant metadata generation approaches, which rely on public platforms and struggle to accommodate governed or large-scale local datasets. The authors propose the first open-source, local-first command-line tool that directly generates Croissant-compliant JSON-LD metadata from local directories via a modular processor registration mechanism, supporting mainstream formats such as Parquet. By eliminating dependence on external platforms, this method significantly enhances the discoverability and reusability of private, high-value datasets. Experimental evaluation across more than 140 datasets—including MIMIC-IV with 886 million rows—demonstrates that the generated metadata achieves 97–100% accuracy, matching or exceeding that of manual curation or standard methods.

0 citationsRead paper

Communicating results in trials with multiple hypotheses or adaptive design features

May 05, 2026

This study addresses the underappreciated challenge of estimation and communication following multiplicity adjustment within the frequentist framework in complex clinical trials, where multiple endpoints, interim data looks, or group comparisons often introduce estimation bias and complicate interpretation, thereby undermining transparency in benefit–risk assessment. By integrating advanced methodologies such as adaptive designs and graphical approaches to multiple testing, the work illustrates through concrete examples the limitations of current strategies in conveying trial results meaningfully. The research underscores the need to critically reevaluate prevailing practices and foster interdisciplinary dialogue to enhance both the accuracy of effect estimation and the clarity of result communication, ultimately informing future methodological standards and regulatory guidance.

0 citationsRead paper