Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究利用大型语言模型和上下文学习解决电子健康记录中机构特定的受保护健康信息脱敏问题,通过不同提示策略提高召回率和精确度。
📝 Abstract
Secondary use of electronic health records requires de-identification, yet existing systems miss \emph{institutionally situated} protected health information (PHI) such as hospital abbreviations, building names, and internal codes whose status is locally determined. We ask whether large language models (LLMs) with in-context learning (ICL) can close this gap and control the precision--recall trade-off. On 100 annotated pediatric oncology notes (5,322 PHI spans) from Texas Children's Hospital, we benchmarked eight LLMs against two purpose-built systems (Stanford TiDE, OpenMed PII) and two pattern-based baselines. Each LLM ran under three prompts of increasing specificity: (1) a HIPAA-aligned baseline, (2) baseline plus the institutional PHI categories it missed, and (3) prompt 2 plus instructions against over-redacting clinical content. We then compared 14~multi-agent and ensemble configurations against the best single prompt, with recall the primary safety metric. LLMs outperformed the purpose-built systems (best F1=0.918$\pm$0.001 vs.\ TiDE 0.779), with advantages concentrated in contextual categories. Naming the missed categories recovered 79\% (48/61) of them, and discouraging over-redaction restored precision. No agentic architecture beat calibrated single-pass prompting (F1 0.906--0.907), but LLM outputs surfaced 414~candidate annotation gaps; re-annotation confirmed 227~PHI spans, against which the final prompt reached recall=0.981 (F1=0.907$\pm$0.002). Well-calibrated ICL resolves both the institutional PHI gap and the precision--recall trade-off in one LLM call per note. LLMs cost more to run than traditional methods, but that cost buys a way to audit the reference standard. LLMs are a legitimate, adaptable alternative to purpose-built de-identification systems; institution-specific prompt development should be the primary adaptation strategy.
Problem

Research questions and friction points this paper is trying to address.

institution-specific PHI
de-identification systems
electronic health records
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-context Learning
Large Language Models
Protected Health Information
Precision-Recall Trade-off
Institution-Specific Prompting
D
Daniel Palacios
Quantitative and Computational Biosciences, Baylor College of Medicine, Houston, Texas, USA; Department of Pediatrics, Baylor College of Medicine, Houston, Texas, USA; Jan and Dan Duncan Neurological Research Institute, Texas Children’s Hospital, Houston, Texas 77030, USA; Data Science Center, Texas Children’s Hospital, Houston, Texas 77030, USA
M
Matthew Brady Neeley
Quantitative and Computational Biosciences, Baylor College of Medicine, Houston, Texas, USA; Department of Pediatrics, Baylor College of Medicine, Houston, Texas, USA; Jan and Dan Duncan Neurological Research Institute, Texas Children’s Hospital, Houston, Texas 77030, USA; Data Science Center, Texas Children’s Hospital, Houston, Texas 77030, USA
A
Angel Adetomike Otto
Section of Hematology-Oncology, Department of Pediatrics, Baylor College of Medicine, Houston, Texas, USA
S
Shalini Dhamodharan
Section of Hematology-Oncology, Department of Pediatrics, Baylor College of Medicine, Houston, Texas, USA
J
John P. Woodhouse
Section of Hematology-Oncology, Department of Pediatrics, Baylor College of Medicine, Houston, Texas, USA
C
Chi-fan Lin
Section of Hematology-Oncology, Department of Pediatrics, Baylor College of Medicine, Houston, Texas, USA
M
Mark Zobeck
Section of Hematology-Oncology, Department of Pediatrics, Baylor College of Medicine, Houston, Texas, USA
Z
Zhandong Liu
Quantitative and Computational Biosciences, Baylor College of Medicine, Houston, Texas, USA; Department of Pediatrics, Baylor College of Medicine, Houston, Texas, USA; Jan and Dan Duncan Neurological Research Institute, Texas Children’s Hospital, Houston, Texas 77030, USA; Data Science Center, Texas Children’s Hospital, Houston, Texas 77030, USA
Hyun-Hwan Jeong
Hyun-Hwan Jeong
Baylor College of Medicine
BioinformaticsComputational BiologyMachine Learning