CTTVAE: Latent Space Structuring for Conditional Tabular Data Generation on Imbalanced Datasets
This work addresses the challenge of generating high-quality synthetic tabular data for severely imbalanced datasets, where existing methods often fail to preserve both fidelity and utility for minority classes in downstream tasks. The authors propose the CTTVAE+TBS framework, which integrates a conditional Transformer-based variational autoencoder with a class-aware triplet boundary loss to restructure the latent space—enhancing intra-class compactness and inter-class separability. Additionally, an adaptive training sampling mechanism dynamically increases minority class exposure during training. Extensive experiments on six real-world datasets demonstrate that the proposed method significantly outperforms baseline approaches, achieving high data fidelity while substantially improving downstream task performance for minority classes—even surpassing models trained on the original imbalanced data—and effectively bridging the privacy gap between interpolation-based and deep generative methods.