π€ AI Summary
Transcriptomic data are often constrained by sample imbalance, technical biases, and ethical regulations, resulting in a scarcity of high-quality real-world datasets. To address this challenge, this work proposes MK-TGAN, a novel generative adversarial network that uniquely integrates a multi-kernel mechanism with graph neural networks to form a knowledge-guided GAN (GNN-GAN) leveraging a gene knowledge graph. By incorporating prior knowledge of geneβgene interactions into the synthetic data generation process, MK-TGAN enhances biological plausibility and utility for downstream tasks. Experimental results demonstrate that MK-TGAN significantly outperforms existing methods, validating the effectiveness and innovation of knowledge-guided strategies in transcriptomic data synthesis.
π Abstract
As biomedical research increasingly relies on data-intensive tools, the quality and utility of datasets are critical. Challenges such as imbalances, biases, and ethical or legal constraints often limit access to high-quality data. Synthetic data generation can help overcome these limitations. Here, we present a comparative analysis of generative models for transcriptomic data, investigating strategies to incorporate prior biological knowledge via gene graphs. This ensures that synthetic data capture real-world gene patterns, maintaining their usefulness for downstream tasks. In particular, we introduce and benchmark three variants of the Generative Adversarial Network. Among the alternatives, MK-TGAN - an innovative multi-kernel, Graph Neural Network-based model - stands out for its performance in terms of both the realism and utility of the generated data. Unlike other methods, MK-TGAN leverages prior knowledge graphs by exploiting graph neural networks. Our results show that prior knowledge integration strategies improve performance, and that MK-TGAN consistently produces synthetic samples with superior realism and biological plausibility.