SNOMED CT Concept Recommendation from Masked Clinical Context
研究解决了SNOMED CT概念推荐难题,特别是在罕见或训练数据中缺失概念的情况下,通过使用掩码临床上下文的方法,并在MIMIC-IV-Note数据上进行了评估。
研究解决了SNOMED CT概念推荐难题,特别是在罕见或训练数据中缺失概念的情况下,通过使用掩码临床上下文的方法,并在MIMIC-IV-Note数据上进行了评估。
This work addresses the challenge of quantifying spatial clustering of probability distributions over graph-structured data by introducing a diffusion distance that incorporates global graph geometry. Built upon a Metropolis–Hastings Markov chain, the proposed measure characterizes the rate at which a distribution converges to uniformity to assess clustering intensity, thereby extending the classical Moran’s I statistic to a global perspective. Theoretically, the distance is linked to spectral graph theory and optimal transport, with formal stability guarantees provided. Empirical evaluations demonstrate superior statistical power over Moran’s I in synthetic benchmarks and reveal nuanced segregation patterns in the distribution of Black populations across 100 U.S. cities—patterns overlooked by traditional local indicators.
Clinical unstructured text contains rich semantic information, yet mining relationships among medical concepts remains limited by the disconnection between co-occurrence statistics and semantic representations. Method: Leveraging SNOMED CT–annotated clinical notes from MIMIC-IV, we systematically analyze the correlation between concept co-occurrence patterns (quantified via normalized pointwise mutual information, NPMI) and semantic similarity derived from pretrained embeddings (ClinicalBERT/BioBERT), revealing only weak correlation—indicating that co-occurrence fails to capture implicit clinical associations. We thus propose a dual-perspective framework integrating co-occurrence and embedding signals: (i) interpretable clinical topics are generated via embedding-based clustering; (ii) clinically meaningful concept pairs—absent in explicit co-occurrence—are identified using embedding proximity. Contribution/Results: The framework significantly improves downstream diagnostic prediction and prognostic modeling (e.g., mortality, readmission). It enhances phenotyping accuracy and annotation completeness, establishing a novel paradigm for clinical decision support.
This study addresses the high cost and poor scalability of manual SNOMED CT concept annotation in clinical text. We propose a lightweight, efficient sequence labeling method that replaces computationally intensive Transformers with a bidirectional GRU architecture, significantly reducing inference overhead while preserving performance. To enhance input representation, we integrate domain-adapted tokenization—combining SpaCy and SciBERT—and incorporate contextual, syntactic, and morphological features, thereby improving robustness against lexical ambiguity and orthographic errors. Evaluated on a MIMIC-IV subset, our model achieves an F1-score of 90%, outperforming conventional rule-based systems and matching state-of-the-art neural models. The approach thus delivers both high accuracy and strong deployability, offering a practical solution for large-scale clinical concept extraction.
This paper addresses the insufficient diversity and suboptimal decision quality of AI agent populations in multi-task environments. To this end, we propose a role-based multi-agent collaboration framework comprising three specialized CNN agents—“Fast-Learning,” “Precision-Learning,” and “Organizational”—built upon VGG16, VGG19, and ResNet50 backbones, respectively, to emulate social collaborative dynamics. We introduce two novel mechanisms: (i) an “endogamous/exogamous” knowledge exchange paradigm, wherein agents exchange knowledge either within or across architectural families, and (ii) a familial model evolution architecture, guided by genetic algorithms for agent optimization and weight-fusion-based knowledge transfer. The framework significantly enhances population robustness and collective decision-making capability. Empirical evaluation shows that descendant agents achieve F1 scores of 82%–95% across diverse classification and prediction tasks, validating the effectiveness and novelty of the role-based design, cross-architectural mating mechanism, and familial agent community structure.
研究解决了SNOMED CT概念推荐难题,特别是在罕见或训练数据中缺失概念的情况下,通过使用掩码临床上下文的方法,并在MIMIC-IV-Note数据上进行了评估。
This work addresses the challenge of quantifying spatial clustering of probability distributions over graph-structured data by introducing a diffusion distance that incorporates global graph geometry. Built upon a Metropolis–Hastings Markov chain, the proposed measure characterizes the rate at which a distribution converges to uniformity to assess clustering intensity, thereby extending the classical Moran’s I statistic to a global perspective. Theoretically, the distance is linked to spectral graph theory and optimal transport, with formal stability guarantees provided. Empirical evaluations demonstrate superior statistical power over Moran’s I in synthetic benchmarks and reveal nuanced segregation patterns in the distribution of Black populations across 100 U.S. cities—patterns overlooked by traditional local indicators.
Clinical unstructured text contains rich semantic information, yet mining relationships among medical concepts remains limited by the disconnection between co-occurrence statistics and semantic representations. Method: Leveraging SNOMED CT–annotated clinical notes from MIMIC-IV, we systematically analyze the correlation between concept co-occurrence patterns (quantified via normalized pointwise mutual information, NPMI) and semantic similarity derived from pretrained embeddings (ClinicalBERT/BioBERT), revealing only weak correlation—indicating that co-occurrence fails to capture implicit clinical associations. We thus propose a dual-perspective framework integrating co-occurrence and embedding signals: (i) interpretable clinical topics are generated via embedding-based clustering; (ii) clinically meaningful concept pairs—absent in explicit co-occurrence—are identified using embedding proximity. Contribution/Results: The framework significantly improves downstream diagnostic prediction and prognostic modeling (e.g., mortality, readmission). It enhances phenotyping accuracy and annotation completeness, establishing a novel paradigm for clinical decision support.
This study addresses the high cost and poor scalability of manual SNOMED CT concept annotation in clinical text. We propose a lightweight, efficient sequence labeling method that replaces computationally intensive Transformers with a bidirectional GRU architecture, significantly reducing inference overhead while preserving performance. To enhance input representation, we integrate domain-adapted tokenization—combining SpaCy and SciBERT—and incorporate contextual, syntactic, and morphological features, thereby improving robustness against lexical ambiguity and orthographic errors. Evaluated on a MIMIC-IV subset, our model achieves an F1-score of 90%, outperforming conventional rule-based systems and matching state-of-the-art neural models. The approach thus delivers both high accuracy and strong deployability, offering a practical solution for large-scale clinical concept extraction.
This paper addresses the insufficient diversity and suboptimal decision quality of AI agent populations in multi-task environments. To this end, we propose a role-based multi-agent collaboration framework comprising three specialized CNN agents—“Fast-Learning,” “Precision-Learning,” and “Organizational”—built upon VGG16, VGG19, and ResNet50 backbones, respectively, to emulate social collaborative dynamics. We introduce two novel mechanisms: (i) an “endogamous/exogamous” knowledge exchange paradigm, wherein agents exchange knowledge either within or across architectural families, and (ii) a familial model evolution architecture, guided by genetic algorithms for agent optimization and weight-fusion-based knowledge transfer. The framework significantly enhances population robustness and collective decision-making capability. Empirical evaluation shows that descendant agents achieve F1 scores of 82%–95% across diverse classification and prediction tasks, validating the effectiveness and novelty of the role-based design, cross-architectural mating mechanism, and familial agent community structure.