Institution profile

Aindo

Industry research
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

ReMIA: a Powerful and Efficient Alternative to Membership Inference Attacks against Synthetic Data Generators

May 14, 2026

This work addresses the vulnerability of synthetic data generators to membership inference attacks, noting that existing approaches incur substantial computational overhead and rely heavily on auxiliary data. To overcome these limitations, the authors propose ReMIA, an efficient and practical privacy risk assessment method that eliminates the need for shadow models. ReMIA requires only two rounds of generator training and no more auxiliary data than the size of the original training set, achieving high attack sensitivity through relative source discrimination rather than absolute membership prediction. Evaluated across multiple tabular datasets and generative models, a classifier-based implementation of ReMIA matches the attack performance of state-of-the-art methods while significantly reducing resource demands, further demonstrating that synthetic data offers a superior privacy-utility trade-off compared to traditional anonymization techniques.

0 citationsRead paper

Graph Conditional Flow Matching for Relational Data Generation

May 21, 2025

Existing multi-table data generation methods struggle to model long-range dependencies and complex foreign-key structures—such as multi-parent tables and heterogeneous associations. To address this, we propose a graph-conditional flow matching framework, the first to introduce flow matching into relational data synthesis. Our approach employs a graph neural network (GNN) to encode the foreign-key dependency graph and guides a record-level denoising neural network, enabling joint, database-wide generation. Crucially, it supports dynamic, cross-table information propagation within arbitrarily connected components, balancing structural flexibility with rich semantic expressiveness. Evaluated on multiple benchmark datasets, our method achieves state-of-the-art synthetic fidelity, significantly improving consistency in multi-table association patterns and joint statistical distributions compared to prior work.

0 citationsRead paper

Contrastive Learning-Based privacy metrics in Tabular Synthetic Datasets

Feb 19, 2025

Existing privacy evaluation methods for synthetic tabular data lack interpretable and quantifiable privacy guarantees. Method: This paper proposes an embedding-based privacy metric grounded in contrastive learning, which uniformly maps heterogeneous attributes into a measurable embedding space—enabling intuitive distance computation and modeling of privacy attacks such as membership inference. Contribution/Results: To our knowledge, this is the first work to apply contrastive learning to quantitative privacy assessment of synthetic data, effectively addressing the challenge of adapting to diverse attribute types (e.g., categorical, numerical, ordinal). The proposed lightweight metric requires no explicit modeling of GDPR-compliance constraints, yet achieves evaluation performance comparable to complex regulatory-compliance models across multiple public benchmarks. It offers significant advantages in computational efficiency and deployment simplicity while preserving rigorous privacy quantification.

0 citationsRead paper
Recent publications

Latest Papers

ReMIA: a Powerful and Efficient Alternative to Membership Inference Attacks against Synthetic Data Generators

May 14, 2026

This work addresses the vulnerability of synthetic data generators to membership inference attacks, noting that existing approaches incur substantial computational overhead and rely heavily on auxiliary data. To overcome these limitations, the authors propose ReMIA, an efficient and practical privacy risk assessment method that eliminates the need for shadow models. ReMIA requires only two rounds of generator training and no more auxiliary data than the size of the original training set, achieving high attack sensitivity through relative source discrimination rather than absolute membership prediction. Evaluated across multiple tabular datasets and generative models, a classifier-based implementation of ReMIA matches the attack performance of state-of-the-art methods while significantly reducing resource demands, further demonstrating that synthetic data offers a superior privacy-utility trade-off compared to traditional anonymization techniques.

0 citationsRead paper

Graph Conditional Flow Matching for Relational Data Generation

May 21, 2025

Existing multi-table data generation methods struggle to model long-range dependencies and complex foreign-key structures—such as multi-parent tables and heterogeneous associations. To address this, we propose a graph-conditional flow matching framework, the first to introduce flow matching into relational data synthesis. Our approach employs a graph neural network (GNN) to encode the foreign-key dependency graph and guides a record-level denoising neural network, enabling joint, database-wide generation. Crucially, it supports dynamic, cross-table information propagation within arbitrarily connected components, balancing structural flexibility with rich semantic expressiveness. Evaluated on multiple benchmark datasets, our method achieves state-of-the-art synthetic fidelity, significantly improving consistency in multi-table association patterns and joint statistical distributions compared to prior work.

0 citationsRead paper

Contrastive Learning-Based privacy metrics in Tabular Synthetic Datasets

Feb 19, 2025

Existing privacy evaluation methods for synthetic tabular data lack interpretable and quantifiable privacy guarantees. Method: This paper proposes an embedding-based privacy metric grounded in contrastive learning, which uniformly maps heterogeneous attributes into a measurable embedding space—enabling intuitive distance computation and modeling of privacy attacks such as membership inference. Contribution/Results: To our knowledge, this is the first work to apply contrastive learning to quantitative privacy assessment of synthetic data, effectively addressing the challenge of adapting to diverse attribute types (e.g., categorical, numerical, ordinal). The proposed lightweight metric requires no explicit modeling of GDPR-compliance constraints, yet achieves evaluation performance comparable to complex regulatory-compliance models across multiple public benchmarks. It offers significant advantages in computational efficiency and deployment simplicity while preserving rigorous privacy quantification.

0 citationsRead paper