Developing synthetic microdata through machine learning for firm-level business surveys
Traditional anonymization of enterprise-level commercial survey data faces significant re-identification risks and struggles to balance confidentiality with analytical utility. To address this, we propose a generative machine learning–based method for synthesizing microdata. Our approach integrates multidimensional distribution matching across geographic and industrial dimensions, employs domain-specific quality metrics, and enforces statistical moment constraints to ensure high-fidelity synthetic data that closely replicates key statistical properties and economic inference outcomes of the original dataset. We successfully generated a synthetic dataset for the 2007 Business Owner Survey and fully reproduced an empirical study published in *Small Business Economics*, thereby validating the method’s statistical validity and reproducibility. This work bridges a critical technical gap in secure commercial survey data dissemination and provides a scalable, regulatory-compliant framework for sharing sensitive microdata.