Dissecting Neuro-Symbolic Quality Assurance for Synthetic Oncology Data Generation

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过神经符号管道生成合成肿瘤数据,解决数据稀缺问题,并通过控制实验验证了各质量保证组件的作用,发现模式验证是关键过滤器。
📝 Abstract
Synthetic clinical data generation with large language models addresses the scarcity that limits cancer staging research, but oncology hallucinations are categorically harmful: one clinically impossible staging assignment contaminates every downstream model trained on it. Neuro-symbolic pipelines validate during generation, yet the contribution of individual quality-assurance components remains unclear. We report three controlled studies isolating gate necessity, constraint attribution, and retrieval conditionality, holding generation protocol, diversity thresholds, and fine-tuning hyperparameters constant across adapter conditions. The symbolic gate enforces schema completeness, ontology coverage against the Systematized Nomenclature of Medicine, and staging-logic consistency under American Joint Committee on Cancer eighth-edition rules. Ungated, 29.9% of records contain schema failures and 20.1% contain clinically invalid staging. Schema validation is the load-bearing filter: within the fully gated corpus it rejects 148 of 512 records, ontology grounding a further 24, and staging-logic validation none---the only generator producing logic violations is already excluded on schema, making clinical-logic validation a generator-conditional safeguard rather than the dominant filter. Retrieval augmentation is strongly model-dependent: it improves gate compliance for one generator by 12.5 percentage points, has no measurable effect for a second, and collapses output in a third. Across gated configurations ontology density is largely unchanged, indicating that symbolic validation improves clinical validity rather than vocabulary richness. Symbolic gating therefore buys corpus validity but no commensurate gain on real lung-cancer notes in this study; retrieval should be evaluated per model, and ontology density should not be reported as a proxy for corpus quality.
Problem

Research questions and friction points this paper is trying to address.

Synthetic clinical data
oncology hallucinations
neuro-symbolic pipelines
quality assurance
staging-logic consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

neuro-symbolic pipeline
quality assurance components
schema validation
retrieval augmentation
ontology coverage
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Laxmigayathri Challa
Department of Information Science, University of North Texas, Denton, TX 76203 USA
Yuhan Zhou
Yuhan Zhou
Ph.D. student, University of North Texas
Data QualityHealth InformaticsData Science
A
Ana Cleveland
Department of Information Science, University of North Texas, Denton, TX 76203 USA
H
Haihua Chen
The Anuradha and Vikas Sinha Department of Data Science, University of North Texas, Denton, TX 76203 USA