Institution profile

University of Göttingen

Academic institutioneurope · de
Official website
Research library174linked papers
Opportunities0open roles
Selected work

Representative Papers

Paraphrase Types Elicit Prompt Engineering Capabilities

Jun 28, 2024Conference on Empirical Methods in Natural Language Processing

This study investigates how linguistic dimensions of prompt formulation—morphology, syntax, and lexis—affect large language model (LLM) task performance. Method: We conduct controlled experiments across 120 diverse tasks using five LLMs, rigorously matching confounding factors such as prompt length and lexical diversity. This enables the first fine-grained, linguistically grounded attribution analysis of prompt rewriting. Contribution/Results: Semantic-preserving rewrites at morphological and lexical levels yield the largest performance gains, revealing LLMs’ high sensitivity to surface-form variation. We propose a multidimensional prompt rewriting generation framework that integrates cross-model behavioral comparison and median gain statistics. Evaluated on Mixtral 8x7B and LLaMA-3-8B, it achieves median task performance improvements of +6.7% and +5.5%, respectively—demonstrating that linguistically informed rewriting systematically enhances prompt robustness and effectiveness.

2 citationsRead paper

Bayesian Penalized Transformation Models: Structured Additive Location-Scale Regression for Arbitrary Conditional Distributions

Apr 11, 2024

This paper addresses the challenge of balancing modeling flexibility and parameter interpretability in conditional distribution regression. We propose the Bayesian Penalized Transformation Model (BPTM), which maps the response variable to a standard reference distribution via a monotone-increasing spline-based transformation function, jointly modeling location, scale, and shape parameters. A key innovation is the incorporation of a smooth Gaussian process prior directly into the transformation function, thereby unifying conditional transformation models and generalized additive distribution models within a single coherent framework. Full Bayesian inference—implemented via the No-U-Turn Sampler (NUTS)—supports structured additive predictors, enabling flexible distributional estimation and natural uncertainty quantification for covariate effects. The method demonstrates robustness and flexibility in extensive simulations and real-data applications, including the Dutch Fourth Growth Study and the Framingham Heart Study. An open-source Python library is provided to facilitate end-to-end distributional regression modeling.

2 citationsRead paper

Belief Updating and Delegation in Multi-Task Human-AI Interaction: Evidence from Controlled Simulations

Feb 02, 2026

This study addresses the limited understanding of how users form and update beliefs about AI across multiple tasks and use these beliefs to decide whether to delegate tasks, particularly when task-specific AI reliability varies substantially. Through a preregistered controlled simulation experiment involving grammar checking, travel planning, and visual question answering, the research employs Bayesian belief modeling and a binary delegation decision paradigm to demonstrate that users’ beliefs about general-purpose AI exhibit path dependence and cross-task transfer. Findings reveal that users do not reset beliefs between tasks; belief updates align directionally with Bayesian predictions but occur at a conservative rate (approximately 50%); and delegation decisions are primarily driven by subjective accuracy beliefs, with self-confidence playing only a secondary inhibitory role. These results challenge the conventional assumption of task independence and offer new foundations for designing multi-task AI systems.

1 citationsRead paper

When Life Gives You AI, Will You Turn It Into A Market for Lemons? Understanding How Information Asymmetries About AI System Capabilities Affect Market Outcomes and Adoption

Jan 29, 2026

This study addresses the severe information asymmetry in the AI consumer market, which impedes users’ ability to identify low-quality systems, suppresses adoption, and undermines market efficiency. Through a simulated market experiment, the research systematically manipulates both the prevalence of low-quality AI systems and the depth of information disclosure, integrating behavioral experiments with Bayesian decision modeling to examine how information asymmetry affects user adoption decisions. The findings provide the first experimental evidence that moderate partial disclosure of system quality effectively mitigates the “lemons problem” in AI markets, significantly improving user decision quality and overall market efficiency. These results offer both theoretical grounding and practical guidance for designing effective information disclosure mechanisms in AI product markets.

1 citationsRead paper

LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification

Aug 12, 2026

This work addresses the limitations of existing financial text classification methods, which often neglect market context and struggle to accurately discern hawkish, dovish, or neutral stances in Federal Reserve communications. To overcome this, the authors propose LabelFusion-TS, a novel system that, for the first time, incorporates financial market time series as an auxiliary modality. The approach fuses a fine-tuned RoBERTa model, prompt-driven large language models, and a time series Transformer, employing a two-stage training strategy to mitigate the scarcity of labeled data. Evaluated on a test set spanning 2015–2022, the model achieves a weighted F1 score of 70.2% using only 240 manually annotated samples—significantly outperforming zero-shot large language models (64.1%)—thereby demonstrating the efficacy of multimodal fusion and few-shot learning in this domain.

0 citationsRead paper
Recent publications

Latest Papers

LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification

Aug 12, 2026

This work addresses the limitations of existing financial text classification methods, which often neglect market context and struggle to accurately discern hawkish, dovish, or neutral stances in Federal Reserve communications. To overcome this, the authors propose LabelFusion-TS, a novel system that, for the first time, incorporates financial market time series as an auxiliary modality. The approach fuses a fine-tuned RoBERTa model, prompt-driven large language models, and a time series Transformer, employing a two-stage training strategy to mitigate the scarcity of labeled data. Evaluated on a test set spanning 2015–2022, the model achieves a weighted F1 score of 70.2% using only 240 manually annotated samples—significantly outperforming zero-shot large language models (64.1%)—thereby demonstrating the efficacy of multimodal fusion and few-shot learning in this domain.

0 citationsRead paper

Risky Business: Measuring The Faithfulness-Safety Tension

Aug 04, 2026

This study addresses the inherent tension between high faithfulness (i.e., monitorability of reasoning traces) and high safety (i.e., rejection of hazardous reasoning) in large reasoning models. To systematically evaluate and intervene in reasoning chains without relying on prompt injection, the authors introduce the HazMart benchmark and propose a Targeted Reasoning Replacement (TRR) method. Through mechanistic interpretability analysis and representation manipulation, they uncover— for the first time—orthogonal internal representation directions governing faithfulness and safety, enabling their independent control. Empirical results reveal a stark trade-off: DeepSeek-R1-Llama-70B achieves 97.5% faithfulness but only 12.3% safety, whereas QwQ-32B attains 73.9% safety at the cost of reduced faithfulness (74.7%). Crucially, representation manipulation improves safety by 9 percentage points without compromising core model capabilities.

0 citationsRead paper

A Physics-Flavored Transformer Network for Parametrizing Contraction Dynamics of Engineered Skeletal Muscle Tissues

Aug 04, 2026

This study addresses the limitations of conventional approaches to functional characterization of engineered skeletal muscle (ESM) tissues, which often rely on oversimplified metrics that neglect critical dynamic features, while mechanistic models remain too complex for scalable application. To bridge this gap, the authors propose a hybrid CNN-Transformer architecture that incorporates stretch-exponential physical priors, thereby embedding biophysical constraints directly into the Transformer for the first time. This enables automatic extraction of high-fidelity dynamic parameters from force–time curves. Leveraging a hybrid training strategy—combining synthetic data pretraining with unsupervised self-alignment on real experimental data—the method achieves efficient and scalable phenotypic analysis under limited-sample conditions. It accurately links idealized models with noisy experimental measurements across multiple cell lines, including a Duchenne muscular dystrophy model, facilitating high-throughput biophysical investigation.

0 citationsRead paper

Network Information Enhances Unreliable News Domain Detection

Aug 03, 2026

This study addresses the growing challenge of fake news detection, exacerbated by generative AI and low-credibility sources mimicking legitimate media. The authors propose a novel paradigm that eschews direct content analysis, instead constructing a domain co-occurrence network from URL-sharing behaviors in Telegram chats. They uncover, for the first time, a homophily effect with respect to source reliability within this network and demonstrate that propagation topology alone can effectively assess domain credibility. Their approach integrates GraphSAGE, multilingual text embeddings, and propagation dynamics to perform reliability classification at the domain level. Experimental results show that the method achieves an accuracy of 0.53 without content features and 0.63 when content is included, yielding a relative improvement of 13–14% over non-graph baselines.

0 citationsRead paper

Nonfundamentalness or missing information ? Evidence from causal-noncausal VARs in macro-finance

Jul 30, 2026

This study investigates whether the noncausal dynamics observed in macroeconomic VAR models stem from genuine non-fundamentalness or from omitted common information that is available to economic agents but unobserved by econometricians. To address this, the paper proposes a hybrid causal–noncausal VARX framework integrated with factor filtering and employs the generalized covariance (GCov) estimator to effectively identify and correct noncausal components. Empirical application to the Stock–Watson monetary policy SVAR demonstrates that the proposed approach substantially attenuates spurious noncausal signals, yielding impulse responses that align more closely with theoretical priors and notably alleviating the “price puzzle.” This refinement enables a more accurate recovery of the underlying causal structure of the economy.

0 citationsRead paper