Institution profile

Leiden University

Academic institutioneurope · nl
Official website
Research library448linked papers
Opportunities0open roles
Selected work

Representative Papers

Undesirable Memorization in Large Language Models: A Survey

Oct 03, 2024arXiv.org

Large language models (LLMs) exhibit “undesired memorization”—the excessive retention and leakage of sensitive training data fragments—posing critical risks for privacy violations and membership inference attacks. Method: We propose the first three-dimensional taxonomy (granularity, retrievability, extractability) to systematically characterize this phenomenon; establish a privacy–utility trade-off analysis framework unifying exposure, membership inference, and related metrics; and extend the study to emerging paradigms including retrieval-augmented generation (RAG) and diffusion language models. Through systematic literature review, empirical attribution analysis, and defense evaluation, we construct a structured knowledge graph and an open-source, dynamically updated literature repository. Results: Our work identifies six key frontiers in LLM memorization governance, advancing the field from ad hoc practice toward rigorous, systematized science.

4 citations1 influentialRead paper

Computational Techniques Enabling the Perception of Virtual Images Exclusive to the Retinal Afterimage

Sep 13, 2022Big Data and Cognitive Computing

How can computational techniques enable perception of virtual images exclusively through retinal afterimages? Method: We propose and implement the first afterimage-specific display method: real-time gaze localization via eye-tracking and visual fixation modeling, coupled with temporally encoded light stimulation to precisely control afterimage generation while strictly isolating the feedforward visual pathway—ensuring image recognition occurs solely within the afterimage. A closed-loop induction–validation experimental system was developed and deployed in real time on AR/VR platforms. Contribution/Results: Participants achieved 92.3% ± 2.1% accuracy in identifying afterimage-specific shapes, demonstrating that afterimages constitute a viable, independent, and controllable information channel. This work establishes retinal afterimages as a programmable visual medium—the first such demonstration—opening new paradigms for neurointerfaces and immersive artistic expression.

2 citationsRead paper

Performance of models for monitoring sustainable development goals from remote sensing: A three-level meta-regression

Jan 07, 2026arXiv.org

This study addresses the lack of consistent evaluation metrics and poor comparability in existing machine learning research for monitoring the United Nations Sustainable Development Goals (SDGs) using remote sensing data. Applying the PRISMA protocol, the authors systematically analyzed 86 experiments from 20 studies through a three-level random-effects meta-regression with double arcsine transformation. The overall mean accuracy was found to be 0.90 [0.86, 0.92], yet model performance exhibited high sensitivity to class imbalance. Notably, 64% of the variance in performance stemmed from between-study differences, with the proportion of the majority class alone accounting for 61% of this heterogeneity. To enhance cross-study comparability and advance standardized evaluation practices in SDG monitoring, the authors advocate for the mandatory reporting of confusion matrices in future studies.

1 citationsRead paper

A Knowledge-Based Language Model: Deducing Grammatical Knowledge in a Multi-Agent Language Acquisition Simulation

Dec 01, 2025

This study addresses the problem of unsupervised acquisition and explicit modeling of linguistic syntactic knowledge. To this end, we propose MODOMA, a multi-agent simulation framework that emulates adult–child interactions to enable data-free language acquisition. We further introduce the Knowledge-based Language Model (KLM), which explicitly represents and parametrically learns grammatical categories—such as functional vs. lexical classes—as interpretable, manipulable structured knowledge. Our method integrates statistical induction with rule-guided learning, enabling fully controllable experiments and faithful replication of human-like developmental patterns. Empirical results demonstrate that child agents consistently induce grammatical categories across varying sample sizes; the KLM exhibits both fluent novel sentence generation and accurate structural parsing. Collectively, this work establishes a novel paradigm for building interpretable, evolvable language models grounded in cognitively plausible mechanisms.

1 citationsRead paper

The anonymization problem in social networks

Sep 24, 2024arXiv.org

This paper addresses the *k*-anonymization problem on social network graphs, aiming to maximize the number of nodes satisfying the *k*-anonymity condition—i.e., each such node must have at least *k*−1 structurally equivalent peers—via structural modifications, primarily edge deletions. We introduce and systematically formulate three novel optimization variants: full *k*-anonymization, partial *k*-anonymization, and budget-constrained *k*-anonymization. Methodologically, we propose a structural-uniqueness-driven edge deletion strategy, surpassing conventional heuristics, and rigorously analyze how anonymity metric selection critically governs the privacy–utility trade-off. Leveraging structural equivalence, we design a reusable computational framework integrating four new heuristic algorithms. Experiments demonstrate that our optimal algorithm retains, on average, 14× more edges than baselines under full *k*-anonymization, and yields 4.8× more *k*-anonymous nodes under budget constraints—significantly improving the balance between privacy protection and graph utility.

1 citationsRead paper
Recent publications

Latest Papers