Responsible Integration of AI in Cancer Genomics: Barriers, Risks, and Pathways to Trustworthy Clinical Translation
本文探讨了AI和NLP在癌症基因组学中的应用挑战,包括证据不一致、可解释性等,并提出通过验证、不确定性感知方法等途径解决这些问题以促进临床转化。
本文探讨了AI和NLP在癌症基因组学中的应用挑战,包括证据不一致、可解释性等,并提出通过验证、不确定性感知方法等途径解决这些问题以促进临床转化。
This study addresses the limitation of existing AI systems that quantify uncertainty without explicit supervisory response logic. We propose a novel "Uncertainty-Action Binding" framework alongside ActionCue visualization technology to bridge this gap. By integrating prioritization strategies with context-aware safety correctors, our approach explicitly combines multiple uncertainty conditions into auditable supervisory decisions, thereby rendering the transition from uncertainty perception to action triggering transparent. Empirical validation across three domains, including healthcare, demonstrates that this method effectively mitigates the black-box decision-making problem. Consequently, it significantly enhances both interpretability and safety in uncertainty management within human-AI collaboration, establishing a verifiable link between model confidence and operational responses.
This work addresses the lack of a unified evaluation framework for DNA data storage encoding schemes, which hinders fair performance comparisons. To this end, we present an open-source, modular benchmarking platform that establishes the first standardized, multidimensional, and extensible evaluation framework specifically designed for DNA storage codecs. Built upon consensus standards from the DNA Data Storage Alliance, the platform integrates diverse benchmark datasets and an automated evaluation pipeline to enable reproducible assessment across key dimensions—including throughput, computational efficiency, error-correction capability (covering substitution, insertion, and deletion errors), and cost. Experimental results demonstrate that no single algorithm dominates across all metrics. The platform supports plug-and-play integration of new algorithms and, through visual and tabular reporting, quantitatively reveals trade-offs among information density, success rate, runtime, and cost, thereby providing empirical guidance for real-world deployment and future format design.
This work addresses the limitation of conventional loss functions—such as cross-entropy—in neglecting structural relationships among classes, which hinders their ability to handle structured label noise or incorporate prior knowledge. The authors propose Conveyance, a novel framework that introduces, for the first time, a unified loss function capable of jointly addressing tasks with structured label spaces, including hierarchical classification, ordinal regression, and multiple-instance learning. By modeling class relationships through a graph structure, the method avoids the need for complex joint distributions or manually designed utility matrices. It further incorporates a double-margin maximization mechanism to optimize decision margins across varying class partitions. The proposed loss enjoys favorable theoretical properties, such as monotonicity and partial convexity, and achieves performance on par with or superior to task-specific methods across multiple benchmark datasets, demonstrating its generality and effectiveness.
This study addresses the lack of standardized evaluation protocols in existing methods for synthesizing health tabular data. To this end, it systematically assesses the performance of seven prominent generative models across four health datasets of varying scales, employing consistent hyperparameter tuning and joint distribution fidelity metrics to ensure a fair comparison. The work introduces a novel, unified evaluation framework that integrates multidimensional quantitative metrics with visual analytics, complemented by domain-informed medical interpretation. Through this approach, the study uncovers critical limitations of current models in adhering to clinical constraints and provides a reproducible, interpretable foundation for selecting appropriate synthetic data generators in healthcare applications.
本文探讨了AI和NLP在癌症基因组学中的应用挑战,包括证据不一致、可解释性等,并提出通过验证、不确定性感知方法等途径解决这些问题以促进临床转化。
This study addresses the limitation of existing AI systems that quantify uncertainty without explicit supervisory response logic. We propose a novel "Uncertainty-Action Binding" framework alongside ActionCue visualization technology to bridge this gap. By integrating prioritization strategies with context-aware safety correctors, our approach explicitly combines multiple uncertainty conditions into auditable supervisory decisions, thereby rendering the transition from uncertainty perception to action triggering transparent. Empirical validation across three domains, including healthcare, demonstrates that this method effectively mitigates the black-box decision-making problem. Consequently, it significantly enhances both interpretability and safety in uncertainty management within human-AI collaboration, establishing a verifiable link between model confidence and operational responses.
This work addresses the lack of a unified evaluation framework for DNA data storage encoding schemes, which hinders fair performance comparisons. To this end, we present an open-source, modular benchmarking platform that establishes the first standardized, multidimensional, and extensible evaluation framework specifically designed for DNA storage codecs. Built upon consensus standards from the DNA Data Storage Alliance, the platform integrates diverse benchmark datasets and an automated evaluation pipeline to enable reproducible assessment across key dimensions—including throughput, computational efficiency, error-correction capability (covering substitution, insertion, and deletion errors), and cost. Experimental results demonstrate that no single algorithm dominates across all metrics. The platform supports plug-and-play integration of new algorithms and, through visual and tabular reporting, quantitatively reveals trade-offs among information density, success rate, runtime, and cost, thereby providing empirical guidance for real-world deployment and future format design.
This work addresses the limitation of conventional loss functions—such as cross-entropy—in neglecting structural relationships among classes, which hinders their ability to handle structured label noise or incorporate prior knowledge. The authors propose Conveyance, a novel framework that introduces, for the first time, a unified loss function capable of jointly addressing tasks with structured label spaces, including hierarchical classification, ordinal regression, and multiple-instance learning. By modeling class relationships through a graph structure, the method avoids the need for complex joint distributions or manually designed utility matrices. It further incorporates a double-margin maximization mechanism to optimize decision margins across varying class partitions. The proposed loss enjoys favorable theoretical properties, such as monotonicity and partial convexity, and achieves performance on par with or superior to task-specific methods across multiple benchmark datasets, demonstrating its generality and effectiveness.
This study addresses the lack of standardized evaluation protocols in existing methods for synthesizing health tabular data. To this end, it systematically assesses the performance of seven prominent generative models across four health datasets of varying scales, employing consistent hyperparameter tuning and joint distribution fidelity metrics to ensure a fair comparison. The work introduces a novel, unified evaluation framework that integrates multidimensional quantitative metrics with visual analytics, complemented by domain-informed medical interpretation. Through this approach, the study uncovers critical limitations of current models in adhering to clinical constraints and provides a reproducible, interpretable foundation for selecting appropriate synthetic data generators in healthcare applications.