Graph Conditional Flow Matching for Relational Data Generation

📅 2025-05-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing multi-table data generation methods struggle to model long-range dependencies and complex foreign-key structures—such as multi-parent tables and heterogeneous associations. To address this, we propose a graph-conditional flow matching framework, the first to introduce flow matching into relational data synthesis. Our approach employs a graph neural network (GNN) to encode the foreign-key dependency graph and guides a record-level denoising neural network, enabling joint, database-wide generation. Crucially, it supports dynamic, cross-table information propagation within arbitrarily connected components, balancing structural flexibility with rich semantic expressiveness. Evaluated on multiple benchmark datasets, our method achieves state-of-the-art synthetic fidelity, significantly improving consistency in multi-table association patterns and joint statistical distributions compared to prior work.

Technology Category

Application Category

📝 Abstract
Data synthesis is gaining momentum as a privacy-enhancing technology. While single-table tabular data generation has seen considerable progress, current methods for multi-table data often lack the flexibility and expressiveness needed to capture complex relational structures. In particular, they struggle with long-range dependencies and complex foreign-key relationships, such as tables with multiple parent tables or multiple types of links between the same pair of tables. We propose a generative model for relational data that generates the content of a relational dataset given the graph formed by the foreign-key relationships. We do this by learning a deep generative model of the content of the whole relational database by flow matching, where the neural network trained to denoise records leverages a graph neural network to obtain information from connected records. Our method is flexible, as it can support relational datasets with complex structures, and expressive, as the generation of each record can be influenced by any other record within the same connected component. We evaluate our method on several benchmark datasets and show that it achieves state-of-the-art performance in terms of synthetic data fidelity.
Problem

Research questions and friction points this paper is trying to address.

Generating multi-table data with complex relational structures
Handling long-range dependencies and foreign-key relationships
Improving synthetic data fidelity for relational datasets
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graph neural network for relational data generation
Flow matching for deep generative modeling
Handling complex foreign-key relationships flexibly
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Davide Scassola
AILAB, University of Trieste, Trieste, Italy
S
Sebastiano Saccani
Aindo, AREA Science Park, Padriciano 99 (TS), Italy
Luca Bortolussi
Luca Bortolussi
Università di Trieste
modelling and simulationexplainable artificial intelligencemachine learningformal methodscyber-physical systems