Institution profile

Insight Centre for Data Analytics

Academic institutioneurope · ie
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

SHAPCA: Consistent and Interpretable Explanations for Machine Learning Models on Spectroscopy Data

Mar 19, 2026

This study addresses the instability in model interpretability caused by high-dimensional and highly collinear spectral data, which often obscures the relationship between original signals and predictive rationale. To resolve this, the authors propose a novel approach that integrates principal component analysis (PCA) with SHAP (SHapley Additive exPlanations). The method first employs PCA for dimensionality reduction to enhance model stability, then maps SHAP-based explanations back to the original spectral space, thereby preserving both global and local interpretability. This work is the first to achieve consistent representation of feature importance—derived from reduced dimensions—in the original input space, effectively mitigating the interpretability disconnect commonly induced by conventional dimensionality reduction. Experimental results demonstrate that the proposed method significantly improves explanation stability across repeated runs and accurately links critical spectral bands to their corresponding biochemical components, thereby enhancing model trustworthiness and practical utility.

0 citationsRead paper

When retrieval outperforms generation: Dense evidence retrieval for scalable fake news detection

Nov 06, 2025

To address the proliferation of misinformation, existing large language model (LLM)-based fact-checking approaches suffer from high computational overhead, severe hallucination risks, and poor deployability. This paper proposes DeReC, a lightweight and efficient fact verification framework that pioneers the integration of general-purpose text embeddings with dense retrieval—replacing LLM-based generative reasoning—and introduces a dedicated classifier for end-to-end verification. By preserving semantic understanding while eliminating autoregressive generation, DeReC significantly reduces computational cost. Experiments show that DeReC achieves an F1 score of 65.58% on RAWFC, outperforming the state-of-the-art L-Defense (61.20%) and accelerating inference by 20× (95% runtime reduction); on LIAR-RAW, it achieves 12× speedup (92% reduction). This work is the first to empirically validate the superiority of non-generative dense retrieval for fact-checking, establishing a new paradigm for scalable, low-cost, and robust verification systems.

0 citationsRead paper

Adapting Psycholinguistic Research for LLMs: Gender-inclusive Language in a Coreference Context

Feb 18, 2025

This study investigates whether large language models (LLMs) exhibit gender neutrality in coreference resolution involving gender-inclusive language (e.g., non-binary or gender-neutral expressions) and uncovers latent cross-linguistic gender biases. Methodologically, it innovatively adapts a French psycholinguistic paradigm—first extended to English and German—combining prompt engineering with controlled cloze tasks across Llama, GPT, and Claude models, validated via statistical significance testing. Results reveal that while English LLMs generally preserve antecedent gender, they exhibit an underlying male bias; in German, this bias is markedly stronger, systematically overriding diverse gender-neutral strategies. Crucially, the study demonstrates how grammatical gender systems amplify implicit biases in LLMs—a previously undocumented phenomenon. It thus establishes a novel methodology and empirical benchmark for cross-linguistic fairness evaluation in NLP.

0 citationsRead paper
Recent publications

Latest Papers

SHAPCA: Consistent and Interpretable Explanations for Machine Learning Models on Spectroscopy Data

Mar 19, 2026

This study addresses the instability in model interpretability caused by high-dimensional and highly collinear spectral data, which often obscures the relationship between original signals and predictive rationale. To resolve this, the authors propose a novel approach that integrates principal component analysis (PCA) with SHAP (SHapley Additive exPlanations). The method first employs PCA for dimensionality reduction to enhance model stability, then maps SHAP-based explanations back to the original spectral space, thereby preserving both global and local interpretability. This work is the first to achieve consistent representation of feature importance—derived from reduced dimensions—in the original input space, effectively mitigating the interpretability disconnect commonly induced by conventional dimensionality reduction. Experimental results demonstrate that the proposed method significantly improves explanation stability across repeated runs and accurately links critical spectral bands to their corresponding biochemical components, thereby enhancing model trustworthiness and practical utility.

0 citationsRead paper

When retrieval outperforms generation: Dense evidence retrieval for scalable fake news detection

Nov 06, 2025

To address the proliferation of misinformation, existing large language model (LLM)-based fact-checking approaches suffer from high computational overhead, severe hallucination risks, and poor deployability. This paper proposes DeReC, a lightweight and efficient fact verification framework that pioneers the integration of general-purpose text embeddings with dense retrieval—replacing LLM-based generative reasoning—and introduces a dedicated classifier for end-to-end verification. By preserving semantic understanding while eliminating autoregressive generation, DeReC significantly reduces computational cost. Experiments show that DeReC achieves an F1 score of 65.58% on RAWFC, outperforming the state-of-the-art L-Defense (61.20%) and accelerating inference by 20× (95% runtime reduction); on LIAR-RAW, it achieves 12× speedup (92% reduction). This work is the first to empirically validate the superiority of non-generative dense retrieval for fact-checking, establishing a new paradigm for scalable, low-cost, and robust verification systems.

0 citationsRead paper

Adapting Psycholinguistic Research for LLMs: Gender-inclusive Language in a Coreference Context

Feb 18, 2025

This study investigates whether large language models (LLMs) exhibit gender neutrality in coreference resolution involving gender-inclusive language (e.g., non-binary or gender-neutral expressions) and uncovers latent cross-linguistic gender biases. Methodologically, it innovatively adapts a French psycholinguistic paradigm—first extended to English and German—combining prompt engineering with controlled cloze tasks across Llama, GPT, and Claude models, validated via statistical significance testing. Results reveal that while English LLMs generally preserve antecedent gender, they exhibit an underlying male bias; in German, this bias is markedly stronger, systematically overriding diverse gender-neutral strategies. Crucially, the study demonstrates how grammatical gender systems amplify implicit biases in LLMs—a previously undocumented phenomenon. It thus establishes a novel methodology and empirical benchmark for cross-linguistic fairness evaluation in NLP.

0 citationsRead paper