Institution profile

UnitedHealth Group

Industry researchnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Semantic Analysis of SNOMED CT Concept Co-occurrences in Clinical Documentation using MIMIC-IV

Sep 03, 2025

Clinical unstructured text contains rich semantic information, yet mining relationships among medical concepts remains limited by the disconnection between co-occurrence statistics and semantic representations. Method: Leveraging SNOMED CT–annotated clinical notes from MIMIC-IV, we systematically analyze the correlation between concept co-occurrence patterns (quantified via normalized pointwise mutual information, NPMI) and semantic similarity derived from pretrained embeddings (ClinicalBERT/BioBERT), revealing only weak correlation—indicating that co-occurrence fails to capture implicit clinical associations. We thus propose a dual-perspective framework integrating co-occurrence and embedding signals: (i) interpretable clinical topics are generated via embedding-based clustering; (ii) clinically meaningful concept pairs—absent in explicit co-occurrence—are identified using embedding proximity. Contribution/Results: The framework significantly improves downstream diagnostic prediction and prognostic modeling (e.g., mortality, readmission). It enhances phenotyping accuracy and annotation completeness, establishing a novel paradigm for clinical decision support.

0 citationsRead paper

Automated SNOMED CT Concept Annotation in Clinical Text Using Bi-GRU Neural Networks

Aug 04, 2025

This study addresses the high cost and poor scalability of manual SNOMED CT concept annotation in clinical text. We propose a lightweight, efficient sequence labeling method that replaces computationally intensive Transformers with a bidirectional GRU architecture, significantly reducing inference overhead while preserving performance. To enhance input representation, we integrate domain-adapted tokenization—combining SpaCy and SciBERT—and incorporate contextual, syntactic, and morphological features, thereby improving robustness against lexical ambiguity and orthographic errors. Evaluated on a MIMIC-IV subset, our model achieves an F1-score of 90%, outperforming conventional rule-based systems and matching state-of-the-art neural models. The approach thus delivers both high accuracy and strong deployability, offering a practical solution for large-scale clinical concept extraction.

0 citationsRead paper

Balanced Area Deprivation Index (bADI): Enhancing social determinants of health indices to strengthen their association with healthcare clinical outcomes, utilization and costs

Jun 09, 2025

Existing Area Deprivation Index (ADI) variants over-rely on housing-related variables—particularly home values—leading to distorted deprivation assessments in high-cost regions and obscuring genuine health inequities and cost heterogeneity. Method: We propose the standardized, balanced ADI (bADI), which mitigates housing-price bias through variable rebalancing and z-score standardization across domains. Leveraging large-scale real-world Medicare data (Fee-for-Service and Medicare Advantage), we employ multivariate modeling and weighted analyses to evaluate bADI’s performance. Contribution/Results: bADI significantly outperforms conventional ADI in predicting clinical outcomes, life expectancy, healthcare utilization, and expenditures. Critically, it more accurately captures cost stratification patterns aligned with health equity principles, thereby enhancing social risk detection and enabling more equitable resource allocation.

0 citationsRead paper
Recent publications

Latest Papers

Semantic Analysis of SNOMED CT Concept Co-occurrences in Clinical Documentation using MIMIC-IV

Sep 03, 2025

Clinical unstructured text contains rich semantic information, yet mining relationships among medical concepts remains limited by the disconnection between co-occurrence statistics and semantic representations. Method: Leveraging SNOMED CT–annotated clinical notes from MIMIC-IV, we systematically analyze the correlation between concept co-occurrence patterns (quantified via normalized pointwise mutual information, NPMI) and semantic similarity derived from pretrained embeddings (ClinicalBERT/BioBERT), revealing only weak correlation—indicating that co-occurrence fails to capture implicit clinical associations. We thus propose a dual-perspective framework integrating co-occurrence and embedding signals: (i) interpretable clinical topics are generated via embedding-based clustering; (ii) clinically meaningful concept pairs—absent in explicit co-occurrence—are identified using embedding proximity. Contribution/Results: The framework significantly improves downstream diagnostic prediction and prognostic modeling (e.g., mortality, readmission). It enhances phenotyping accuracy and annotation completeness, establishing a novel paradigm for clinical decision support.

0 citationsRead paper

Automated SNOMED CT Concept Annotation in Clinical Text Using Bi-GRU Neural Networks

Aug 04, 2025

This study addresses the high cost and poor scalability of manual SNOMED CT concept annotation in clinical text. We propose a lightweight, efficient sequence labeling method that replaces computationally intensive Transformers with a bidirectional GRU architecture, significantly reducing inference overhead while preserving performance. To enhance input representation, we integrate domain-adapted tokenization—combining SpaCy and SciBERT—and incorporate contextual, syntactic, and morphological features, thereby improving robustness against lexical ambiguity and orthographic errors. Evaluated on a MIMIC-IV subset, our model achieves an F1-score of 90%, outperforming conventional rule-based systems and matching state-of-the-art neural models. The approach thus delivers both high accuracy and strong deployability, offering a practical solution for large-scale clinical concept extraction.

0 citationsRead paper

Balanced Area Deprivation Index (bADI): Enhancing social determinants of health indices to strengthen their association with healthcare clinical outcomes, utilization and costs

Jun 09, 2025

Existing Area Deprivation Index (ADI) variants over-rely on housing-related variables—particularly home values—leading to distorted deprivation assessments in high-cost regions and obscuring genuine health inequities and cost heterogeneity. Method: We propose the standardized, balanced ADI (bADI), which mitigates housing-price bias through variable rebalancing and z-score standardization across domains. Leveraging large-scale real-world Medicare data (Fee-for-Service and Medicare Advantage), we employ multivariate modeling and weighted analyses to evaluate bADI’s performance. Contribution/Results: bADI significantly outperforms conventional ADI in predicting clinical outcomes, life expectancy, healthcare utilization, and expenditures. Critically, it more accurately captures cost stratification patterns aligned with health equity principles, thereby enhancing social risk detection and enabling more equitable resource allocation.

0 citationsRead paper