Institution profile

TH Köln

Academic institutioneurope · de
Official website
Research library20linked papers
Opportunities0open roles
Selected work

Representative Papers

Perception-Aware Bias Detection for Query Suggestions

Jan 07, 2026International Workshop on Algorithmic Bias in Search and Recommendation

This work addresses the challenge of effectively detecting systemic thematic bias in query suggestion systems, which is hindered by data sparsity, insufficient contextual metadata, and the transient nature of user perception. Building upon the bias detection framework introduced by Bonart et al., the authors propose a novel “perception-aware” bias metric grounded in principles from perceptual psychology, specifically tailored for person-centric search scenarios. By explicitly modeling biases that are actually noticeable to users, the proposed approach substantially enhances the real-world relevance of bias detection outcomes. Experimental validation demonstrates that this refined pipeline more accurately identifies systemically embedded biases that users can perceive, thereby improving both the practical utility and interpretability of bias assessments in search systems.

3 citationsRead paper

Validating Search Query Simulations: A Taxonomy of Measures

Jan 16, 2026

This study addresses the lack of standardized validation criteria for user simulators in information retrieval evaluation, which undermines the reliability of simulation outcomes. Through a systematic literature review, it proposes the first structured taxonomy of metrics specifically designed for validating simulated search queries. The work empirically analyzes the interrelationships among these metrics across four diverse datasets and, based on the findings, offers tailored validation recommendations for different application scenarios. To foster standardization and reproducibility in simulation-based evaluation, the authors also release an open-source toolkit implementing commonly used validation metrics, thereby supporting future research extension and benchmarking.

1 citationsRead paper

CIR at iKAT SCAI 2026: Exploring Clarification Need Prediction in Agentic Conversational Search

Jul 22, 2026

This work proposes an end-to-end conversational search agent that integrates query rewriting, retrieval reranking, answer generation, and a clarification mechanism to enhance both interaction efficiency and retrieval quality. The key innovation lies in embedding a neural clarification module—comprising clarification need prediction and clarifying question generation—directly into the conversational search pipeline, enabling proactive clarification of user intent. Experimental results demonstrate that the proposed approach significantly improves overall system performance and validate the effectiveness of various clarification prediction models in real-world conversational scenarios.

0 citationsRead paper

A Novel Method to Evaluate Models on Unreliable, Noisy and Inconsistent Labels: Adaptive Resolution Label Aggregation (ARLA)

Jul 13, 2026

Existing evaluation methods for segmentation models struggle to handle label noise, inconsistency, or unreliability. This work proposes an Adaptive Resolution Label Aggregation (ARLA) mechanism—the first evaluation framework specifically designed for such challenging scenarios. ARLA dynamically aligns and rescales both predictions and labels during inference, leveraging multi-scale aggregation and domain knowledge to adaptively match label fidelity with prevailing noise levels. By doing so, it effectively extracts a clear performance signal from noisy annotations. In flood prediction tasks, the method substantially mitigates issues arising from inconsistent labeling in forested regions and erroneous labels in cloud-covered areas, significantly enhancing evaluation reliability.

0 citationsRead paper

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks

Apr 22, 2026

This study addresses the absence of automated, high-quality frameworks for generating and evaluating actuarial reasoning tasks, which hinders alignment with international actuarial education standards and effective assessment of large language models’ (LLMs) domain-specific capabilities. To bridge this gap, the authors propose the first multi-agent collaborative LLM pipeline, featuring a role-separated adapter architecture, a first-order bounded repair loop, and a Wikipedia-augmented module to automatically generate and validate both multiple-choice and open-ended questions conforming to the International Actuarial Association syllabus. An online evaluation platform is also developed. Experiments yield 200 high-quality questions and benchmark 50 mainstream LLMs, demonstrating significantly improved question quality through the proposed method. Notably, low-cost open-weight models exhibit strong performance, and substantial discrepancies are observed between multiple-choice assessments and LLM-as-Judge scoring in evaluating model competence.

0 citationsRead paper
Recent publications

Latest Papers

CIR at iKAT SCAI 2026: Exploring Clarification Need Prediction in Agentic Conversational Search

Jul 22, 2026

This work proposes an end-to-end conversational search agent that integrates query rewriting, retrieval reranking, answer generation, and a clarification mechanism to enhance both interaction efficiency and retrieval quality. The key innovation lies in embedding a neural clarification module—comprising clarification need prediction and clarifying question generation—directly into the conversational search pipeline, enabling proactive clarification of user intent. Experimental results demonstrate that the proposed approach significantly improves overall system performance and validate the effectiveness of various clarification prediction models in real-world conversational scenarios.

0 citationsRead paper

A Novel Method to Evaluate Models on Unreliable, Noisy and Inconsistent Labels: Adaptive Resolution Label Aggregation (ARLA)

Jul 13, 2026

Existing evaluation methods for segmentation models struggle to handle label noise, inconsistency, or unreliability. This work proposes an Adaptive Resolution Label Aggregation (ARLA) mechanism—the first evaluation framework specifically designed for such challenging scenarios. ARLA dynamically aligns and rescales both predictions and labels during inference, leveraging multi-scale aggregation and domain knowledge to adaptively match label fidelity with prevailing noise levels. By doing so, it effectively extracts a clear performance signal from noisy annotations. In flood prediction tasks, the method substantially mitigates issues arising from inconsistent labeling in forested regions and erroneous labels in cloud-covered areas, significantly enhancing evaluation reliability.

0 citationsRead paper

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks

Apr 22, 2026

This study addresses the absence of automated, high-quality frameworks for generating and evaluating actuarial reasoning tasks, which hinders alignment with international actuarial education standards and effective assessment of large language models’ (LLMs) domain-specific capabilities. To bridge this gap, the authors propose the first multi-agent collaborative LLM pipeline, featuring a role-separated adapter architecture, a first-order bounded repair loop, and a Wikipedia-augmented module to automatically generate and validate both multiple-choice and open-ended questions conforming to the International Actuarial Association syllabus. An online evaluation platform is also developed. Experiments yield 200 high-quality questions and benchmark 50 mainstream LLMs, demonstrating significantly improved question quality through the proposed method. Notably, low-cost open-weight models exhibit strong performance, and substantial discrepancies are observed between multiple-choice assessments and LLM-as-Judge scoring in evaluating model competence.

0 citationsRead paper

LLMs taking shortcuts in test generation: A study with SAP HANA and LevelDB

Apr 15, 2026

This study investigates whether large language models (LLMs) rely on shallow heuristics or memorization rather than genuine reasoning when generating software tests, particularly for complex systems absent from their training data. By comparing LLM-generated tests for the open-source LevelDB and the proprietary SAP HANA database, and integrating mutation testing scores, iterative compile-feedback repair loops, and the Mitchell mechanism-focused evaluation framework, this work pioneers the application of mechanism-oriented cognitive science methods to software testing. The findings reveal that while LLMs perform well on familiar systems, their effectiveness degrades significantly on unseen systems, often prioritizing syntactic compilability over semantic correctness. This highlights a lack of robust reasoning capabilities in current LLMs and establishes a novel paradigm for evaluating LLM-based reasoning in software engineering contexts.

0 citationsRead paper

A technical curriculum on language-oriented artificial intelligence in translation and specialised communication

Feb 12, 2026

This study addresses the insufficient AI literacy among language-oriented professionals in translation and specialized communication by designing and implementing the first systematic AI curriculum tailored for non-technical language service practitioners. The course covers foundational concepts including vector embeddings, tokenization, neural network basics, and the Transformer architecture, aiming to cultivate computational thinking, algorithmic awareness, algorithmic agency, and digital resilience. Implemented within a master’s program at TH Köln (Cologne University of Applied Sciences), the curriculum demonstrated pedagogical efficacy, with findings indicating that advanced instructional scaffolding—such as direct instructor support—is essential to further enhance learning outcomes. This work thus offers an innovative paradigm for integrating AI literacy into language service education.

0 citationsRead paper