Institution profile

University of Groningen

Academic institutioneurope · nl
Official website
Research library378linked papers
Opportunities0open roles
Selected work

Representative Papers

Undesirable Memorization in Large Language Models: A Survey

Oct 03, 2024arXiv.org

Large language models (LLMs) exhibit “undesired memorization”—the excessive retention and leakage of sensitive training data fragments—posing critical risks for privacy violations and membership inference attacks. Method: We propose the first three-dimensional taxonomy (granularity, retrievability, extractability) to systematically characterize this phenomenon; establish a privacy–utility trade-off analysis framework unifying exposure, membership inference, and related metrics; and extend the study to emerging paradigms including retrieval-augmented generation (RAG) and diffusion language models. Through systematic literature review, empirical attribution analysis, and defense evaluation, we construct a structured knowledge graph and an open-source, dynamically updated literature repository. Results: Our work identifies six key frontiers in LLM memorization governance, advancing the field from ad hoc practice toward rigorous, systematized science.

4 citations1 influentialRead paper

How disentangled are your classification uncertainties?

Aug 22, 2024arXiv.org

Reliable disentanglement of aleatoric and epistemic uncertainty remains elusive in Bayesian neural networks. Method: We systematically evaluate two mainstream approaches—information-theoretic decomposition and Gaussian logits—by proposing the first quantitative experimental criterion for measuring disentanglement quality, and constructing a reproducible diagnostic framework that employs uncertainty-sensitive experiments to assess independence between the two uncertainty types. Contribution/Results: Both methods exhibit significant cross-contamination; while the information-theoretic approach performs comparatively better, neither achieves reliable separation, as substantial entanglement persists. This work establishes the first benchmarking framework for evaluating uncertainty disentanglement performance, revealing fundamental limitations of current methodologies. It provides both theoretical insights and empirical standards to guide future research on principled uncertainty decomposition.

4 citationsRead paper

LongSpeech: A Scalable Benchmark for Transcription, Translation and Understanding in Long Speech

Jan 20, 2026

This work addresses the limited robustness and integrative reasoning capabilities of current speech models in long-form scenarios—such as meeting transcription and spoken document understanding—despite their strong performance on short utterances. To bridge this gap, we introduce LongSpeech, the first large-scale, extensible multitask benchmark for long speech, comprising over 100,000 audio segments averaging ten minutes each. LongSpeech supports diverse tasks including automatic speech recognition, speech translation, summarization, language identification, speaker counting, content disentanglement, and question answering. The benchmark is built from heterogeneous data sources and features multidimensional manual and automatic annotations, standardized evaluation protocols, and a reproducible construction pipeline, offering a unified platform for long-speech research. Preliminary evaluations reveal substantial performance gaps in state-of-the-art models, particularly in cross-task generalization and higher-order reasoning.

1 citations1 influentialRead paper
Recent publications

Latest Papers