Institution profile

University of Trier

Academic institutioneurope · de
Official website
Research library14linked papers
Opportunities0open roles
Selected work

Representative Papers

TriQua: Reconciling Granularity and Context in Factuality Evaluation

Aug 05, 2026

This work addresses the trade-off between granularity and context in factuality evaluation of large language models: atomic facts often lack contextual nuance, while broad statements resist fine-grained assessment. To reconcile this, the authors propose TriQua, a framework that adaptively models factual claims according to their complexity—representing simple assertions as standard triples and encoding complex ones as hyper-relational facts enriched with contextual qualifiers. This approach preserves atomicity while retaining essential context, enabling interpretable, fine-grained error localization. The framework further introduces TriQuaScore, a metric for quantifying factuality at the level of structured factual units. Experiments demonstrate that TriQua achieves robust performance in decomposition quality and alignment with human annotations, outperforming existing methods on evidence-based fact verification tasks.

0 citationsRead paper

A JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs

Jul 14, 2026

This work addresses the memory bottleneck of key-value (KV) cache in large language models during long-context inference by introducing FlashJoLT, the first framework to model the KV cache as a third-order tensor. FlashJoLT integrates partial Tucker decomposition to compress token and feature dimensions, Johnson–Lindenstrauss random projections combined with low-bit quantization to preserve residual information, and a Lagrangian dual formulation to jointly optimize compression ratio and reconstruction accuracy. Randomized SVD further accelerates tensor decomposition. Evaluated on Mistral-7B and LLaMA-2-13B, FlashJoLT achieves 2–3× KV cache compression with relative Frobenius errors as low as 0.006–0.009, while maintaining perplexity, GSM8K accuracy, and RULER retrieval performance on par with baselines, and accelerating decomposition by 5–13×.

0 citationsRead paper

2-Head 2D Returning Finite Automata

Jun 25, 2026

This study investigates the computational power of two-head finite automata in two-dimensional picture language recognition and their relationship with context-free matrix grammars (CFMGs) and returning pushdown automata (RPDAs). It introduces two novel models of two-head returning finite automata operating on rectangular pictures: the 2-HRFA, which allows backward moves, and the B2-HRFA, which enforces synchronous head movement—a constraint newly introduced in this work. Through formal language-theoretic analysis and closure properties, the paper establishes that the class of languages recognized by 2-HRFA is a proper subset of those recognized by RPDAs and incomparable with CFMGs. Furthermore, it demonstrates that B2-HRFA strictly lies between RFA and 2-HRFA in recognition power, thereby establishing a strict hierarchy among these three models.

0 citationsRead paper

The Unsampled Truth: Psychometrics in SLMs Measure Prompt Artifacts, Not Psychological Constructs

Jun 02, 2026

This study addresses whether the outputs of small language models (SLMs) in psychometric tasks stem from genuine semantic reasoning or are primarily driven by artifacts of prompt formulation. The authors propose the first diagnostic framework capable of disentangling the influence of such prompt artifacts, systematically manipulating role framing, instructions, item content, and option labels while employing controlled experiments and variance decomposition techniques to quantify the relative contributions of semantic signals versus prompt-induced artifacts. Findings reveal that prompt artifacts frequently dominate model responses, substantially undermining their psychometric validity. The proposed framework not only effectively identifies these confounding influences but also offers a novel pathway for evaluating and enhancing the semantic comprehension capabilities of large language models.

0 citationsRead paper
Recent publications

Latest Papers

TriQua: Reconciling Granularity and Context in Factuality Evaluation

Aug 05, 2026

This work addresses the trade-off between granularity and context in factuality evaluation of large language models: atomic facts often lack contextual nuance, while broad statements resist fine-grained assessment. To reconcile this, the authors propose TriQua, a framework that adaptively models factual claims according to their complexity—representing simple assertions as standard triples and encoding complex ones as hyper-relational facts enriched with contextual qualifiers. This approach preserves atomicity while retaining essential context, enabling interpretable, fine-grained error localization. The framework further introduces TriQuaScore, a metric for quantifying factuality at the level of structured factual units. Experiments demonstrate that TriQua achieves robust performance in decomposition quality and alignment with human annotations, outperforming existing methods on evidence-based fact verification tasks.

0 citationsRead paper

A JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs

Jul 14, 2026

This work addresses the memory bottleneck of key-value (KV) cache in large language models during long-context inference by introducing FlashJoLT, the first framework to model the KV cache as a third-order tensor. FlashJoLT integrates partial Tucker decomposition to compress token and feature dimensions, Johnson–Lindenstrauss random projections combined with low-bit quantization to preserve residual information, and a Lagrangian dual formulation to jointly optimize compression ratio and reconstruction accuracy. Randomized SVD further accelerates tensor decomposition. Evaluated on Mistral-7B and LLaMA-2-13B, FlashJoLT achieves 2–3× KV cache compression with relative Frobenius errors as low as 0.006–0.009, while maintaining perplexity, GSM8K accuracy, and RULER retrieval performance on par with baselines, and accelerating decomposition by 5–13×.

0 citationsRead paper

2-Head 2D Returning Finite Automata

Jun 25, 2026

This study investigates the computational power of two-head finite automata in two-dimensional picture language recognition and their relationship with context-free matrix grammars (CFMGs) and returning pushdown automata (RPDAs). It introduces two novel models of two-head returning finite automata operating on rectangular pictures: the 2-HRFA, which allows backward moves, and the B2-HRFA, which enforces synchronous head movement—a constraint newly introduced in this work. Through formal language-theoretic analysis and closure properties, the paper establishes that the class of languages recognized by 2-HRFA is a proper subset of those recognized by RPDAs and incomparable with CFMGs. Furthermore, it demonstrates that B2-HRFA strictly lies between RFA and 2-HRFA in recognition power, thereby establishing a strict hierarchy among these three models.

0 citationsRead paper

The Unsampled Truth: Psychometrics in SLMs Measure Prompt Artifacts, Not Psychological Constructs

Jun 02, 2026

This study addresses whether the outputs of small language models (SLMs) in psychometric tasks stem from genuine semantic reasoning or are primarily driven by artifacts of prompt formulation. The authors propose the first diagnostic framework capable of disentangling the influence of such prompt artifacts, systematically manipulating role framing, instructions, item content, and option labels while employing controlled experiments and variance decomposition techniques to quantify the relative contributions of semantic signals versus prompt-induced artifacts. Findings reveal that prompt artifacts frequently dominate model responses, substantially undermining their psychometric validity. The proposed framework not only effectively identifies these confounding influences but also offers a novel pathway for evaluating and enhancing the semantic comprehension capabilities of large language models.

0 citationsRead paper