Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建临床案例基准,评估未来信息暴露对大型语言模型判断的影响,采用时间屏蔽方法减少后见之明偏差。
📝 Abstract
Clinical decisions are prospective, but clinical language models are often evaluated on retrospective records that reveal the final diagnosis, treatment response, and outcome. Such evaluations may reward the use of future information rather than reasoning under the uncertainty present at the decision point. We introduce a paired benchmark for measuring outcome-conditioned shifts consistent with hindsight bias in clinical temporal reasoning. It contains 171 case reports from the PubMed Central Open Access Subset---40 sepsis and 131 GLP-1/diabetes cases---represented as both textual narratives and human-annotated and LLM-generated textual time series (TTS). For each case, questions are tied to a clinically meaningful cutoff and paired with a prospective reference answer and an outcome-consistent \emph{hindsight trap}. Models answer each question using either a TTS truncated at the cutoff or the complete timeline; additional conditions vary the narrative source (original or synthetic) and TTS annotation source (human or LLM). We evaluate accuracy (Acc), hindsight trap rate (HTR), answer instability rate (AIR), and hindsight bias rate (HBR), each of which captures different signals of hindsight bias. Across GPT 5.6 Sol, Gemma 4, GLM 5.2, and Opus 5, full timeline exposure produces consistent hindsight-sensitive shifts, while temporal masking reduces bias without lowering accuracy.
Problem

Research questions and friction points this paper is trying to address.

Clinical Decisions
Retrospective Records
Hindsight Bias
Temporal Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

paired benchmark
outcome-conditioned shifts
hindsight bias
temporal masking
🔎 Similar Papers
No similar papers found.
M
Misaki Matsuura
Case School of Engineering, Case Western Reserve University, USA
Sayantan Kumar
Sayantan Kumar
Postdoc, National Library of Medicine, National Institutes on Health
Machine LearningAI in healthcarebiomedical informaticsprecision medicine
O
Ojas Kadam
Rice University, USA
J
Jeremy C. Weiss
National Library of Medicine, National Institutes of Health, USA