Institution profile

Children’s Hospital of Philadelphia

Academic institutionnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Extracting Post-Acute Sequelae of SARS-CoV-2 Infection Symptoms from Clinical Notes via Hybrid Natural Language Processing

Aug 17, 2025

Post-acute sequelae of SARS-CoV-2 infection (PASC) exhibit high symptom heterogeneity and temporal dynamics, hindering accurate clinical identification from unstructured electronic health records. Method: We developed an end-to-end hybrid NLP pipeline integrating rule-based named entity recognition with a fine-tuned BERT model for assertion classification, augmented by clinical text normalization and a curated PASC-specific terminology dictionary. Results: The system achieved an F1-score of 0.82 in single-center validation and 0.76 in ten-center external validation, with an average processing time of 2.45 seconds per note; assertion outputs showed strong correlation with ground-truth annotations (Spearman ρ > 0.83, *P* < 0.0001). Its key contribution is the first synergistic integration of structured linguistic rules and deep learning–based assertion modeling for PASC symptom extraction—enhancing both cross-center generalizability and clinical interpretability, thereby enabling robust large-scale PASC epidemiological studies.

0 citationsRead paper
Recent publications

Latest Papers

Extracting Post-Acute Sequelae of SARS-CoV-2 Infection Symptoms from Clinical Notes via Hybrid Natural Language Processing

Aug 17, 2025

Post-acute sequelae of SARS-CoV-2 infection (PASC) exhibit high symptom heterogeneity and temporal dynamics, hindering accurate clinical identification from unstructured electronic health records. Method: We developed an end-to-end hybrid NLP pipeline integrating rule-based named entity recognition with a fine-tuned BERT model for assertion classification, augmented by clinical text normalization and a curated PASC-specific terminology dictionary. Results: The system achieved an F1-score of 0.82 in single-center validation and 0.76 in ten-center external validation, with an average processing time of 2.45 seconds per note; assertion outputs showed strong correlation with ground-truth annotations (Spearman ρ > 0.83, *P* < 0.0001). Its key contribution is the first synergistic integration of structured linguistic rules and deep learning–based assertion modeling for PASC symptom extraction—enhancing both cross-center generalizability and clinical interpretability, thereby enabling robust large-scale PASC epidemiological studies.

0 citationsRead paper