Institution profile

Johns Hopkins University

Academic institutionnorthamerica · us
Official website
Research library1,830linked papers
Opportunities0open roles
Selected work

Representative Papers

Medical Hallucinations in Foundation Models and Their Impact on Healthcare

Feb 26, 2025arXiv.org

Medical foundation models may generate “hallucinations”—factual, logical, or evidence-inconsistent errors—that jeopardize clinical decision-making and patient safety. To address this, we first propose a multidimensional taxonomy of medical hallucinations and establish a real-world, clinician-annotated benchmark dataset derived from authentic clinical cases; we further validate its clinical impact via an international physician survey. Methodologically, we integrate expert annotation, empirical behavioral surveys, and large language model (LLM) evaluation to systematically assess the efficacy of chain-of-thought (CoT) reasoning and retrieval-augmented generation (RAG) in mitigating hallucinations. Results show both techniques significantly reduce hallucination rates, yet residual hallucinations remain clinically hazardous. Building on these findings, we introduce a patient-safety-centered AI governance and ethics framework, offering theoretical foundations and actionable pathways for responsible deployment of medical AI. (149 words)

41 citationsRead paper

Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions

Feb 28, 2024arXiv.org

Current large language models (LLMs) lack rigorous, clinically grounded evaluation of reasoning interpretability in complex medical decision-making. Method: We introduce JAMA Clinical Challenge and Medbullets—two high-difficulty, multiple-choice clinical benchmarks featuring authoritative, fine-grained expert explanations—the first such resources designed for real-world clinical scenarios. Evaluation employs zero-shot and few-shot prompting, automated explanation quality scoring, dual-blinded clinical expert assessment (Cohen’s κ = 0.82), and comparative analysis. Contribution/Results: Seven state-of-the-art LLMs exhibit substantially lower performance on these benchmarks than on conventional exam-style benchmarks. Their generated explanations frequently contain logical gaps and factual hallucinations, exposing critical deficiencies in clinical-grade causal reasoning and domain-specific knowledge integration—highlighting a fundamental gap between current LLM capabilities and safe, interpretable clinical deployment.

14 citations2 influentialRead paper

Exponential Lower Bounds for Locally Decodable and Correctable Codes for Insertions and Deletions

Nov 01, 2021IEEE Annual Symposium on Foundations of Computer Science

This work investigates the existence of locally decodable codes (LDCs) under insertion-deletion (insdel) errors. Addressing a long-standing open conjecture, we prove—*for the first time*—that no 2-query linear insdel LDC exists. Moreover, for any constant query complexity $q geq 3$, we establish an exponential lower bound $exp(Omega(n))$ on the code length, significantly stronger than the polynomial bounds known for Hamming-error LDCs. Our approach constructs a hard insdel error distribution and combines information-theoretic analysis with novel coding reduction techniques. This reveals a fundamental separation between insdel LDCs and Hamming LDCs—a separation that persists even in the adaptive decoding and private-key settings. The results characterize the theoretical limits of local error correction against synchronization errors and provide the first tight lower bounds for insdel coding.

7 citations2 influentialRead paper

Current Agents Fail to Leverage World Model as Tool for Foresight

Jan 07, 2026arXiv.org

Current agents exhibit underutilization, misinterpretation of predictions, or performance degradation when employing generative world models for prospective reasoning. This work presents the first systematic evaluation of vision-language model–based agents’ ability to strategically invoke and integrate world model predictions across multitask scenarios, combining visual question answering with agent benchmark tasks and introducing attribution analysis to quantitatively measure simulation usage. The study reveals that agents proactively invoke simulations in fewer than 1% of cases, approximately 15% of predictions are misused, and enforced simulation use can degrade performance by up to 5%. These findings expose a cognitive bottleneck in how agents interpret and incorporate predictive information, underscoring the urgent need for calibrated mechanisms to enable reliable prospective reasoning.

3 citationsRead paper

How Well Do Supervised 3D Models Transfer to Medical Imaging Tasks?

Jan 20, 2025International Conference on Learning Representations

Limited large-scale annotated data hinders effective pretraining for 3D medical image segmentation. Method: This paper introduces AbdomenAtlas 1.1—a high-quality abdominal CT dataset comprising 9,262 cases, covering 25 anatomical structures and seven tumor classes with pseudo-labels—and proposes an efficient supervised 3D pretraining paradigm based on a 3D U-Net variant. It integrates voxel-level supervision, cross-task transfer, pseudo-label augmentation, and standardized collaborative annotation. Contribution/Results: The first systematic validation demonstrates that supervised 3D pretraining significantly improves downstream segmentation performance. Remarkably, only 21 finely annotated cases suffice to match the performance of unsupervised pretraining on 5,050 cases. Under few-shot settings (21 cases / 672 masks / 40 GPU-hours), our method surpasses existing pretrained models across all metrics and achieves comparable accuracy to large-scale unsupervised approaches (1,152 GPU-hours), substantially enhancing data efficiency and clinical deployability.

3 citationsRead paper
Recent publications

Latest Papers