Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data

πŸ“… 2026-08-13
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Transcriptomic data are often constrained by sample imbalance, technical biases, and ethical regulations, resulting in a scarcity of high-quality real-world datasets. To address this challenge, this work proposes MK-TGAN, a novel generative adversarial network that uniquely integrates a multi-kernel mechanism with graph neural networks to form a knowledge-guided GAN (GNN-GAN) leveraging a gene knowledge graph. By incorporating prior knowledge of gene–gene interactions into the synthetic data generation process, MK-TGAN enhances biological plausibility and utility for downstream tasks. Experimental results demonstrate that MK-TGAN significantly outperforms existing methods, validating the effectiveness and innovation of knowledge-guided strategies in transcriptomic data synthesis.
πŸ“ Abstract
As biomedical research increasingly relies on data-intensive tools, the quality and utility of datasets are critical. Challenges such as imbalances, biases, and ethical or legal constraints often limit access to high-quality data. Synthetic data generation can help overcome these limitations. Here, we present a comparative analysis of generative models for transcriptomic data, investigating strategies to incorporate prior biological knowledge via gene graphs. This ensures that synthetic data capture real-world gene patterns, maintaining their usefulness for downstream tasks. In particular, we introduce and benchmark three variants of the Generative Adversarial Network. Among the alternatives, MK-TGAN - an innovative multi-kernel, Graph Neural Network-based model - stands out for its performance in terms of both the realism and utility of the generated data. Unlike other methods, MK-TGAN leverages prior knowledge graphs by exploiting graph neural networks. Our results show that prior knowledge integration strategies improve performance, and that MK-TGAN consistently produces synthetic samples with superior realism and biological plausibility.
Problem

Research questions and friction points this paper is trying to address.

synthetic transcriptomic data
data imbalance
prior biological knowledge
data utility
ethical constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

knowledge-guided generation
graph neural networks
synthetic transcriptomic data
multi-kernel TGAN
biological knowledge integration
πŸ”Ž Similar Papers
No similar papers found.
F
Francesca Pia Panaccione
DEIB - Dipartimento Elettronica, Informazione e Bioingegneria, Politecnico di Milano, Milan, Italy
S
Sofia Mongardi
DEIB - Dipartimento Elettronica, Informazione e Bioingegneria, Politecnico di Milano, Milan, Italy
Marco Masseroli
Marco Masseroli
Politecnico di Milano
BioinformaticsMachine LearningData Bases
Pietro Pinoli
Pietro Pinoli
Research Fellow, Politecnico di Milano
BioinformaticsMachine LearningBig Data