Institution profile

Martin Luther University Halle-Wittenberg

Academic institutioneurope · de
Official website
Research library9linked papers
Opportunities0open roles
Selected work

Representative Papers

FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings

Aug 07, 2026

This work addresses the challenge of evidence provenance errors in financial question answering, which frequently arise due to similar disclosures across sections, reporting periods, and companies in SEC filings. The authors introduce the first evidence-oriented benchmark for 10-K/10-Q–based financial QA and retrieval, requiring systems to precisely align entities, reporting periods, and contextual cues to locate supporting evidence. They contribute a novel set of human-annotated hard negatives—including confounding passages—and define three distinct evaluation tasks: retrieval, re-ranking, and hard negative discrimination. Evaluated on 1,185 annotated instances, even the strongest 7B-scale model achieves only 44.8% Recall@10 using BM25, instruction-tuned embeddings, and pairwise re-ranking. Performance drops by 13.0–20.5 percentage points when hard negatives are introduced, underscoring the task’s difficulty.

0 citationsRead paper

Ontology-Aware Design Patterns for Clinical AI Systems: Translating Reification Theory into Software Architecture

Apr 02, 2026

Clinical AI systems often suffer from ontological distortions in data due to fragmented medical documentation workflows, billing incentives, and inconsistent terminologies, and current architectures lack effective mechanisms to address these issues. This work proposes seven ontology-aware design patterns—four of which are novel—that operationalize reification theory into software architecture for the first time. These patterns encompass data validation, signal preservation, drift monitoring, dual-ontology representation, feedback disruption mitigation, terminology evolution, and regulatory compliance adaptation. By integrating ontology engineering, terminology governance, and regulatory requirements, the proposed pattern suite establishes an ontologically resilient reference architecture for clinical AI. The paper demonstrates the combined application of these patterns in a diabetes risk prediction scenario, providing a foundation for empirical evaluation.

0 citationsRead paper

Enhanced Mycelium of Thought (EMoT): A Bio-Inspired Hierarchical Reasoning Architecture with Strategic Dormancy and Mnemonic Encoding

Mar 25, 2026

This work addresses the limitations of existing large language model (LLM) reasoning approaches—such as Chain-of-Thought (CoT) and Tree-of-Thought (ToT)—which lack persistent memory, strategic dormancy mechanisms, and cross-domain integration capabilities, thereby struggling with complex, multi-domain problems. Inspired by mycelial networks, the authors propose a novel four-layer hierarchical reasoning architecture (Micro/Meso/Macro/Meta) that uniquely integrates hierarchical topology, strategic thought dormancy and reactivation, and a memory palace framework unifying five types of memory encoding within a single system, evaluated via an LLM-as-Judge paradigm. Experiments demonstrate significant performance gains over CoT on cross-domain composite tasks (4.8 vs. 4.4), with improved stability; however, the method incurs approximately 33× the computational overhead of baseline methods and exhibits over-reasoning on simple problems.

0 citationsRead paper

Antisocial behavior towards large language model users: experimental evidence

Jan 14, 2026

This study investigates whether negative social attitudes toward users of large language models (LLMs) translate into tangible punitive behavior. Employing a two-stage online experiment, participants were given the opportunity to sacrifice their own earnings to reduce the rewards of others based on whether they used or refrained from using an LLM to complete a task, thereby quantifying the intensity of social sanctions. The research provides the first behavioral evidence that, despite enhancing task efficiency, LLM use elicits significant social punishment; moreover, disclosing LLM usage exacerbates credibility biases, leading to misjudgments and excessive penalties. Findings reveal that pure LLM users had, on average, 36% of their earnings actively destroyed, with punishment intensity increasing monotonically with actual usage levels. Individuals who misrepresented their LLM usage faced even harsher sanctions.

0 citationsRead paper

Initial data analysis of the national German transplantation registry with a focus on kidney transplantation

Jan 05, 2026

This study addresses data quality challenges in the German Organ Transplantation Registry (TxReg), where missingness, inconsistencies, and ambiguity in event-time variable selection compromise research reliability. Analyzing data from 14,954 recipients and 9,964 donors between 2006 and 2016, this work systematically characterizes conflicts and complementarities among multi-source variables, identifying 168 cross-verifiable fields. By integrating missingness pattern analysis, decision tree modeling, and multi-source consistency checks, the study delineates the underlying missing data structure and proposes targeted imputation strategies. Findings reveal that while some tables exhibit missing rates exceeding 50%, key variables retain high imputation potential. Moreover, event-time analyses prove highly sensitive to variable selection, underscoring the need for careful curation. This work establishes a robust data foundation for future high-quality research leveraging TxReg.

0 citationsRead paper
Recent publications

Latest Papers

FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings

Aug 07, 2026

This work addresses the challenge of evidence provenance errors in financial question answering, which frequently arise due to similar disclosures across sections, reporting periods, and companies in SEC filings. The authors introduce the first evidence-oriented benchmark for 10-K/10-Q–based financial QA and retrieval, requiring systems to precisely align entities, reporting periods, and contextual cues to locate supporting evidence. They contribute a novel set of human-annotated hard negatives—including confounding passages—and define three distinct evaluation tasks: retrieval, re-ranking, and hard negative discrimination. Evaluated on 1,185 annotated instances, even the strongest 7B-scale model achieves only 44.8% Recall@10 using BM25, instruction-tuned embeddings, and pairwise re-ranking. Performance drops by 13.0–20.5 percentage points when hard negatives are introduced, underscoring the task’s difficulty.

0 citationsRead paper

Ontology-Aware Design Patterns for Clinical AI Systems: Translating Reification Theory into Software Architecture

Apr 02, 2026

Clinical AI systems often suffer from ontological distortions in data due to fragmented medical documentation workflows, billing incentives, and inconsistent terminologies, and current architectures lack effective mechanisms to address these issues. This work proposes seven ontology-aware design patterns—four of which are novel—that operationalize reification theory into software architecture for the first time. These patterns encompass data validation, signal preservation, drift monitoring, dual-ontology representation, feedback disruption mitigation, terminology evolution, and regulatory compliance adaptation. By integrating ontology engineering, terminology governance, and regulatory requirements, the proposed pattern suite establishes an ontologically resilient reference architecture for clinical AI. The paper demonstrates the combined application of these patterns in a diabetes risk prediction scenario, providing a foundation for empirical evaluation.

0 citationsRead paper

Enhanced Mycelium of Thought (EMoT): A Bio-Inspired Hierarchical Reasoning Architecture with Strategic Dormancy and Mnemonic Encoding

Mar 25, 2026

This work addresses the limitations of existing large language model (LLM) reasoning approaches—such as Chain-of-Thought (CoT) and Tree-of-Thought (ToT)—which lack persistent memory, strategic dormancy mechanisms, and cross-domain integration capabilities, thereby struggling with complex, multi-domain problems. Inspired by mycelial networks, the authors propose a novel four-layer hierarchical reasoning architecture (Micro/Meso/Macro/Meta) that uniquely integrates hierarchical topology, strategic thought dormancy and reactivation, and a memory palace framework unifying five types of memory encoding within a single system, evaluated via an LLM-as-Judge paradigm. Experiments demonstrate significant performance gains over CoT on cross-domain composite tasks (4.8 vs. 4.4), with improved stability; however, the method incurs approximately 33× the computational overhead of baseline methods and exhibits over-reasoning on simple problems.

0 citationsRead paper

Antisocial behavior towards large language model users: experimental evidence

Jan 14, 2026

This study investigates whether negative social attitudes toward users of large language models (LLMs) translate into tangible punitive behavior. Employing a two-stage online experiment, participants were given the opportunity to sacrifice their own earnings to reduce the rewards of others based on whether they used or refrained from using an LLM to complete a task, thereby quantifying the intensity of social sanctions. The research provides the first behavioral evidence that, despite enhancing task efficiency, LLM use elicits significant social punishment; moreover, disclosing LLM usage exacerbates credibility biases, leading to misjudgments and excessive penalties. Findings reveal that pure LLM users had, on average, 36% of their earnings actively destroyed, with punishment intensity increasing monotonically with actual usage levels. Individuals who misrepresented their LLM usage faced even harsher sanctions.

0 citationsRead paper

Initial data analysis of the national German transplantation registry with a focus on kidney transplantation

Jan 05, 2026

This study addresses data quality challenges in the German Organ Transplantation Registry (TxReg), where missingness, inconsistencies, and ambiguity in event-time variable selection compromise research reliability. Analyzing data from 14,954 recipients and 9,964 donors between 2006 and 2016, this work systematically characterizes conflicts and complementarities among multi-source variables, identifying 168 cross-verifiable fields. By integrating missingness pattern analysis, decision tree modeling, and multi-source consistency checks, the study delineates the underlying missing data structure and proposes targeted imputation strategies. Findings reveal that while some tables exhibit missing rates exceeding 50%, key variables retain high imputation potential. Moreover, event-time analyses prove highly sensitive to variable selection, underscoring the need for careful curation. This work establishes a robust data foundation for future high-quality research leveraging TxReg.

0 citationsRead paper