Institution profile

RheinMain University of Applied Sciences

Academic institutioneurope · de
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

BAFIS: Dataset + Framework to assess occupational Bias and Human Preference in modern Text-to-image Models

Jun 18, 2026

This study addresses the significant gender and racial biases exhibited by contemporary text-to-image models in generating occupation-related imagery, which conventional evaluation metrics often fail to capture due to their inability to reflect human subjective judgments of fairness. To bridge this gap, the authors propose BAFIS—a fairness evaluation framework that integrates human preference feedback with multilingual prompts—and construct a dataset of 21,140 images aligned with official employment statistics. Using this framework, they systematically assess occupational bias, image quality, and prompt alignment across leading models including Midjourney v6.1, Stable Diffusion 3 Medium, and DALL·E 3. The work pioneers the incorporation of human preference annotations into bias evaluation, revealing systematic disparities in model outputs and demonstrating only partial correlation between human feedback and traditional automated metrics, thereby underscoring the critical role of human judgment in developing equitable text-to-image generation systems.

0 citationsRead paper

Agentic Insight Generation in VSM Simulations

Apr 14, 2026

This work addresses the challenges of extracting actionable insights from complex Value Stream Mapping (VSM) simulations, which are often time-consuming, error-prone, and hindered by difficulties in discerning subtle contextual differences. To overcome these limitations, the authors propose a decoupled two-stage agent architecture that separates orchestration from data analysis and integrates a domain knowledge–driven multi-hop reasoning mechanism. This approach enables precise data source selection and effective capture of contextual nuances within lightweight contexts. Evaluated across multiple state-of-the-art large language models, the method achieves a peak accuracy of 86%, demonstrating strong robustness and generalization capability across diverse data structures.

0 citationsRead paper

Multicalibration for LLM-based Code Generation

Dec 09, 2025

This work addresses the misalignment between confidence scores and actual correctness in code generation by large language models (LLMs). We propose the first multidimensional calibration framework tailored for code generation, introducing multicalibration—previously unexplored in the code domain—to enable conditional calibration across fine-grained attributes such as programming language, problem complexity, and generated code length. We implement four multicalibration methods on state-of-the-art models—including Qwen3 Coder, GPT-OSS, and DeepSeek-R1-Distill—within a function synthesis benchmark. Experimental results show that our approach improves skill score by 1.03 over the uncalibrated baseline and outperforms conventional calibration methods by 0.37. To foster reproducibility and further research, we publicly release the first standardized code calibration dataset, comprising generated code samples, likelihood estimates, and ground-truth correctness labels.

0 citationsRead paper

A parameter study for LLL and BKZ with application to shortest vector problems

Feb 07, 2025

This work investigates the efficiency of solving the shortest vector problem (SVP) derived from Learning With Errors (LWE), to support security analysis of NIST’s 2024 standardized module-lattice key encapsulation mechanism (ML-KEM). Method: Through systematic numerical experiments and statistical evaluation, we quantify— for the first time—the success probability of recovering SVP solutions using the LLL and BKZ lattice reduction algorithms across varying lattice dimensions, moduli, and BKZ block sizes. Contribution/Results: We identify critical parameter sensitivity patterns: BKZ achieves substantially higher solution recovery rates for medium-scale LWE instances when the block size is ≥30; conversely, LLL remains practically effective in low-dimensional and small-modulus settings. Based on these findings, we propose an empirically grounded lattice-basis-reduction efficacy criterion tailored for LWE security assessment. This criterion provides concrete, data-driven guidance for selecting secure and efficient parameters in lattice-based cryptography.

0 citationsRead paper
Recent publications

Latest Papers

BAFIS: Dataset + Framework to assess occupational Bias and Human Preference in modern Text-to-image Models

Jun 18, 2026

This study addresses the significant gender and racial biases exhibited by contemporary text-to-image models in generating occupation-related imagery, which conventional evaluation metrics often fail to capture due to their inability to reflect human subjective judgments of fairness. To bridge this gap, the authors propose BAFIS—a fairness evaluation framework that integrates human preference feedback with multilingual prompts—and construct a dataset of 21,140 images aligned with official employment statistics. Using this framework, they systematically assess occupational bias, image quality, and prompt alignment across leading models including Midjourney v6.1, Stable Diffusion 3 Medium, and DALL·E 3. The work pioneers the incorporation of human preference annotations into bias evaluation, revealing systematic disparities in model outputs and demonstrating only partial correlation between human feedback and traditional automated metrics, thereby underscoring the critical role of human judgment in developing equitable text-to-image generation systems.

0 citationsRead paper

Agentic Insight Generation in VSM Simulations

Apr 14, 2026

This work addresses the challenges of extracting actionable insights from complex Value Stream Mapping (VSM) simulations, which are often time-consuming, error-prone, and hindered by difficulties in discerning subtle contextual differences. To overcome these limitations, the authors propose a decoupled two-stage agent architecture that separates orchestration from data analysis and integrates a domain knowledge–driven multi-hop reasoning mechanism. This approach enables precise data source selection and effective capture of contextual nuances within lightweight contexts. Evaluated across multiple state-of-the-art large language models, the method achieves a peak accuracy of 86%, demonstrating strong robustness and generalization capability across diverse data structures.

0 citationsRead paper

Multicalibration for LLM-based Code Generation

Dec 09, 2025

This work addresses the misalignment between confidence scores and actual correctness in code generation by large language models (LLMs). We propose the first multidimensional calibration framework tailored for code generation, introducing multicalibration—previously unexplored in the code domain—to enable conditional calibration across fine-grained attributes such as programming language, problem complexity, and generated code length. We implement four multicalibration methods on state-of-the-art models—including Qwen3 Coder, GPT-OSS, and DeepSeek-R1-Distill—within a function synthesis benchmark. Experimental results show that our approach improves skill score by 1.03 over the uncalibrated baseline and outperforms conventional calibration methods by 0.37. To foster reproducibility and further research, we publicly release the first standardized code calibration dataset, comprising generated code samples, likelihood estimates, and ground-truth correctness labels.

0 citationsRead paper

A parameter study for LLL and BKZ with application to shortest vector problems

Feb 07, 2025

This work investigates the efficiency of solving the shortest vector problem (SVP) derived from Learning With Errors (LWE), to support security analysis of NIST’s 2024 standardized module-lattice key encapsulation mechanism (ML-KEM). Method: Through systematic numerical experiments and statistical evaluation, we quantify— for the first time—the success probability of recovering SVP solutions using the LLL and BKZ lattice reduction algorithms across varying lattice dimensions, moduli, and BKZ block sizes. Contribution/Results: We identify critical parameter sensitivity patterns: BKZ achieves substantially higher solution recovery rates for medium-scale LWE instances when the block size is ≥30; conversely, LLL remains practically effective in low-dimensional and small-modulus settings. Based on these findings, we propose an empirically grounded lattice-basis-reduction efficacy criterion tailored for LWE security assessment. This criterion provides concrete, data-driven guidance for selecting secure and efficient parameters in lattice-based cryptography.

0 citationsRead paper