Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入“相同输入重运行”方法及六个可靠性指标,评估临床LLM代理在重复运行时的行为一致性问题。
📝 Abstract
A clinical agent benchmark can report the same verdict on identical inputs while the agent files a materially different order on each run. Such agents order tests, request medications and place referrals, yet benchmarks typically score one run per task and rarely ask whether identical inputs produce identical actions; MedAgentBench, the benchmark we use, scores a single attempt and says so. To measure this gap we introduce "same-input rerun", which replays a task with every input held fixed and compares the orders rather than the score, with six reliability metrics, and apply it to 1000 MedAgentBench runs across 50 tasks from its five write-capable families, two open-weight models below ten billion parameters quantised to four bits, and two temperatures. The study establishes that action-level divergence exists and can pass unrecorded by the score, not that any rate generalises. Under the 8B model at temperature 0.7, all 43 ordering groups emit a different set of orders across five identical runs, 26 emit the order on some runs and not others, and 28 record a different coded value, dose or analyte. In 22 of those 43 the benchmark reports the same failing verdict for materially different behaviour, as it does for all 10 divergent groups of the 4B model at 0.7. Orders also reach different endpoints across runs, one of which the record server rejects while the agent is told it succeeded. These findings motivate repeated-run evaluation, action-level stability reporting and execution-faithful environment feedback in clinical-agent benchmarks.
Problem

Research questions and friction points this paper is trying to address.

clinical LLM agents
same-input rerun
action-level reliability
benchmark evaluation
MedAgentBench
Innovation

Methods, ideas, or system contributions that make the work stand out.

same-input rerun
action-level reliability
clinical LLM agents
reliability metrics
execution-faithful feedback
🔎 Similar Papers
No similar papers found.
R
Rohith Reddy Bellibatlu
Florida International University, Miami, FL, USA
M
Manpreet Singh
Boston University, Boston, MA, USA
Z
Zhoutian Han
Stevens Institute of Technology, Hoboken, NJ, USA
Wenbin Zhang
Wenbin Zhang
Florida International University
AI AlignmentGenerative AIHealth CareInterdisciplinary