Institution profile

Tilburg University

Academic institutioneurope · nl
Official website
Research library121linked papers
Opportunities0open roles
Selected work

Representative Papers

Memorization to Generalization: Emergence of Diffusion Models from Associative Memory

May 27, 2025

This work investigates the memory–generalization phase transition in diffusion models under varying training data scales. We propose a *correlational memory* perspective: training corresponds to memory encoding, while generation implements memory retrieval. We establish, for the first time, a theoretical connection between diffusion models and Hopfield networks, deriving necessary and sufficient conditions for the emergence of *spurious attractors*—hallucinated states—at the critical memory load threshold. Leveraging energy landscape analysis, dynamical systems modeling, and empirical validation on DDPM and DDIM, we confirm the universality of this phenomenon. Results show that models operate dominantly in memory mode under small-data regimes, shift toward generalization with large-scale data, and exhibit spurious attractors in the critical regime—unifying explanations for memory overload and implicit manifold learning. This work provides a cross-disciplinary theoretical framework and falsifiable predictions for understanding the intrinsic mechanisms of diffusion models.

3 citationsRead paper

Natural Language Generation

Oct 24, 2018Theoretical Issues In Natural Language Processing

Natural language generation (NLG) lacks a unified conceptual framework and clearly delineated disciplinary boundaries, leading to ambiguity in its scope relative to other NLP subfields such as machine translation and dialogue systems. Method: This paper systematically surveys NLG’s research landscape and historical evolution, focusing on core tasks—including data-to-text generation, text summarization, and image captioning—and proposes a novel taxonomy grounded in the dual primitives of “content selection” and “realization.” It traces methodological shifts from rule-based and statistical approaches to early neural models, analyzing how task-driven evaluation paradigms have evolved. Contribution/Results: The work establishes NLG’s formal disciplinary boundaries for the first time, constructs a widely cited conceptual framework and discipline map, clarifies terminological consensus, and identifies large language models as catalysts for methodological convergence across NLP subfields. These contributions lay foundational theoretical groundwork for NLG research and practice.

2 citations1 influentialRead paper

Measuring Semantic Information Production in Generative Diffusion Models

Jun 12, 2025

This study investigates when semantic category decisions emerge during the reverse denoising process of generative diffusion models. Method: We propose an information flow rate metric based on the time derivative of conditional entropy to quantify the dynamic emergence of semantic information; we design an online Bayesian classifier coupled with a conditional entropy estimation framework, and conduct experiments on DDPM using CIFAR-10 and a 1D Gaussian mixture model. Contribution/Results: We find that semantic information flow peaks in the mid-stage of denoising and vanishes toward the final steps; entropy change rates diverge significantly across classes, revealing temporal heterogeneity in semantic decision-making. This work challenges the conventional assumption of uniform temporal semantics and establishes the first interpretable, computationally tractable framework for analyzing the temporal dynamics of semantic decisions in diffusion models—providing a novel theoretical tool for understanding generative mechanisms and enabling controllable generation.

1 citationsRead paper

A Pairwise Differencing Distribution Regression Approach for Network Models

Aug 05, 2026

This study addresses the incidental parameter problem arising in two-way fixed effects network models when the outcome variable is sparse or observed at extreme quantiles. The authors propose a distributional regression approach based on multi-threshold binarization combined with conditional maximum likelihood estimation, which effectively eliminates fixed effects through pairwise differencing to identify structural parameters. They innovatively derive the joint asymptotic distribution of estimators across different thresholds, enabling the construction of simultaneous confidence bands and the development of cross-threshold equality tests for coefficients. Monte Carlo simulations demonstrate that the method exhibits low bias and accurate coverage even under sparsity. An application to bilateral trade data reveals significant heterogeneity in the impact of key trade barriers across the conditional distribution of trade flows.

0 citationsRead paper

Lean-verified lower bounds for the Shannon capacity of odd cycles

Jul 31, 2026

This work addresses the long-standing open problem of establishing lower bounds on the Shannon capacity of small odd cycles. By integrating Gao’s iterative method with techniques introduced by Itty et al., we present the first complete formal verification of these lower bounds within the Lean theorem prover, ensuring both mathematical rigor and computational precision. Leveraging constructions of independent sets from graph theory together with iterative optimization algorithms, we obtain the strongest known lower bounds to date: notably, Θ(C₇) ≥ 3.2588…, along with improved bounds for odd cycles C₁₁ through C₂₃. These results substantially advance the theoretical understanding of the Shannon capacity for small odd-cycle graphs.

0 citationsRead paper
Recent publications

Latest Papers

A Pairwise Differencing Distribution Regression Approach for Network Models

Aug 05, 2026

This study addresses the incidental parameter problem arising in two-way fixed effects network models when the outcome variable is sparse or observed at extreme quantiles. The authors propose a distributional regression approach based on multi-threshold binarization combined with conditional maximum likelihood estimation, which effectively eliminates fixed effects through pairwise differencing to identify structural parameters. They innovatively derive the joint asymptotic distribution of estimators across different thresholds, enabling the construction of simultaneous confidence bands and the development of cross-threshold equality tests for coefficients. Monte Carlo simulations demonstrate that the method exhibits low bias and accurate coverage even under sparsity. An application to bilateral trade data reveals significant heterogeneity in the impact of key trade barriers across the conditional distribution of trade flows.

0 citationsRead paper

Lean-verified lower bounds for the Shannon capacity of odd cycles

Jul 31, 2026

This work addresses the long-standing open problem of establishing lower bounds on the Shannon capacity of small odd cycles. By integrating Gao’s iterative method with techniques introduced by Itty et al., we present the first complete formal verification of these lower bounds within the Lean theorem prover, ensuring both mathematical rigor and computational precision. Leveraging constructions of independent sets from graph theory together with iterative optimization algorithms, we obtain the strongest known lower bounds to date: notably, Θ(C₇) ≥ 3.2588…, along with improved bounds for odd cycles C₁₁ through C₂₃. These results substantially advance the theoretical understanding of the Shannon capacity for small odd-cycle graphs.

0 citationsRead paper

LLMs for health: Perceived benefits, risks, intention to use AI chatbots, and willingness to self-disclose across sensitive health topics

Jul 10, 2026

This study investigates users’ willingness to use AI chatbots and their self-disclosure of health information across physical and mental health topics varying in sensitivity. Employing a 2 (physical vs. psychological) × 2 (low vs. high sensitivity) mixed online experiment with a representative sample of 1,388 Dutch participants, the research systematically examines how topic sensitivity and individual characteristics jointly shape user decisions—an aspect previously unexplored. Findings reveal that perceived benefits positively predict both usage intention and information disclosure, whereas perceived risks exert a negative influence. Usage intention is significantly higher for low-sensitivity topics, and individual traits notably moderate these relationships. The study offers theoretical grounding and practical implications for the design of health-focused AI systems.

0 citationsRead paper

DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information

Jun 29, 2026

This work addresses the challenge of detecting sensitive personally identifiable information in conversational data from privacy-sensitive domains such as healthcare and social sciences. To support privacy-preserving and responsible data sharing, the authors construct a multilingual synthetic dialogue dataset spanning eight scenarios, nineteen entity types, and eleven languages. Dialogues are generated using large language models, manually validated, synthesized into speech, and transcribed with Whisper to create aligned text–speech pairs. The study innovatively combines automatic annotation projection with human correction to achieve fine-grained cross-modal labeling. The released dataset includes high-quality annotations and a Transformer-based baseline named entity recognition model. Experimental evaluation through inter-annotator agreement, translation quality, and benchmark performance demonstrates the dataset’s validity and practical utility.

0 citationsRead paper

Computing Lewis weights to high precision using local relative smoothness

Jun 28, 2026

This work addresses the problem of efficiently and accurately computing ℓ_p-Lewis weights (for p ≥ 4) of a matrix, which quantify the importance of its rows. By alternating between primal and dual formulations of the underlying optimization problem and integrating leverage score iteration with a locally relative smooth gradient descent method, the authors propose a novel algorithm that significantly reduces computational overhead while maintaining high accuracy. Specifically, the proposed approach improves the iteration complexity from O(p³ log(m/ε)) to O(p² log(m/ε)), achieving the current best-known bound on the number of iterations required for ε-approximate Lewis weight computation.

0 citationsRead paper