Institution profile

Forschungszentrum Jülich

Academic institutioneurope · de
Official website
Research library233linked papers
Opportunities0open roles
Selected work

Representative Papers

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Jan 17, 2026

This work addresses the challenge that existing AI agent benchmarks inadequately evaluate performance on real-world, complex, and long-horizon command-line tasks. To bridge this gap, the authors introduce a novel evaluation benchmark comprising 89 high-difficulty terminal tasks, all derived from authentic workflows and accompanied by isolated execution environments, human-authored reference solutions, and automated verification tests. The benchmark is designed to ensure realism, verifiability, and diversity, substantially narrowing the disparity between practical scenarios and current model evaluation paradigms. Experimental results demonstrate that even state-of-the-art agents achieve success rates below 65% on this benchmark. The paper further provides comprehensive error analysis and publicly releases the dataset and evaluation toolchain to support future research in this domain.

9 citations1 influentialRead paper

MEmilio -- A high performance Modular EpideMIcs simuLatIOn software for multi-scale and comparative simulations of infectious disease dynamics

Feb 11, 2026

This work proposes a unified, modular, high-performance simulation framework to address the fragmentation in current infectious disease modeling ecosystems, which hinders cross-model comparison and deployment across model types, spatial scales, and computational platforms. For the first time, the framework integrates compartmental models, meta-population models, and agent-based models within a single architecture, enabling multi-scale and comparable epidemic dynamics simulations. By standardizing representations of spatial, demographic, and mobility data, coupling a high-performance C++ core with a Python interface, and incorporating uncertainty quantification and parameter inference tools, the framework supports seamless deployment—from laptops to high-performance computing environments—significantly lowering barriers to reuse and accelerating the development of simulation-driven epidemic response capabilities.

2 citationsRead paper

Using street view images and visual LLMs to predict heritage values for governance support: Risks, ethics, and policy implications

Dec 22, 2025arXiv.org

Sweden lacks a national registry of architectural heritage values, which hinders the formulation of effective building retrofit policies. This study addresses this gap by pioneering the integration of multimodal large language models (MLLMs) with street-view imagery to predict heritage values for over 150,000 buildings nationwide—encompassing approximately 5 million square meters of heated floor area—using zero-shot learning. Beyond delivering critical data to inform national retrofitting initiatives, the research systematically examines the ethical implications, transparency, and policy impacts of deploying vision-based large language models in public governance. The findings illuminate both the transformative potential and inherent risks of such AI applications, offering an innovative paradigm for AI-driven cultural heritage management.

2 citationsRead paper

Toward Scalable Normalizing Flows for the Hubbard Model

Jan 26, 2026

This work investigates the effective scaling of normalizing flows to larger lattice sizes and lower temperatures in the Hubbard model for efficient learning of its Boltzmann distribution. Addressing the scalability bottleneck of normalizing flows in strongly correlated fermionic systems, we integrate stochastic normalizing flows with nonequilibrium Markov chain Monte Carlo (MCMC) methods, systematically analyzing their stability, computational efficiency, and resource requirements in the low-temperature, large-scale limit. Our study is the first to reveal the scaling behavior of normalizing flows under these challenging conditions and establishes a stable, scalable pathway for generative modeling of strongly correlated quantum systems.

1 citationsRead paper

Ontology-aligned structuring and reuse of multimodal materials data and workflows towards automatic reproduction

Jan 18, 2026

This work addresses the challenge posed by unstructured textual descriptions of density functional theory (DFT) workflows in computational materials science, which hinder reproducibility and systematic comparison. We propose the first framework that integrates domain ontologies with large language models (LLMs) to automatically extract stacking fault energy computation protocols from scientific literature. Through a multi-stage text filtering pipeline and tailored prompt engineering, our approach aligns extracted workflows with established materials ontologies—including CMSO, ASMO, and PLDO—and constructs a structured atomRDF knowledge graph. This method achieves, for the first time, semantic structuring and cross-study alignment of DFT workflows, substantially enhancing the machine-readability, transparency, and reusability of computational materials data.

1 citationsRead paper
Recent publications

Latest Papers