Toward Workflow-Aware Benchmarking for Healthcare NLP Agents

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对医疗NLP代理在动态工作流环境下的评估不足问题,提出了一种新的分阶段评估协议,涵盖文档更新、证据检索、患者消息传递和分流交接四个任务模板。
📝 Abstract
Large language model (LLM) agents are increasingly proposed for healthcare tasks such as clinical documentation, evidence retrieval, patient messaging, and care coordination. Yet many evaluations remain limited to static medical question answering or one-shot generation, under-representing longitudinal state, interruptions, and human handoffs. We introduce an episode-level evaluation protocol for healthcare NLP agents. The protocol separates evidence across model, agent, and simulated-workflow behavior; specifies a five-field episode schema; and defines annotation and scoring for state continuity, evidence traceability, and escalation decisions. It is instantiated as four task templates: documentation update, evidence retrieval, patient messaging, and triage handoff. The protocol does not claim to measure clinical outcomes or deployment value. Instead, it supplies a reproducible intermediate evaluation layer between static benchmarks and prospective workflow studies, with an explicit cost-sensitive treatment of missed versus unnecessary escalation.
Problem

Research questions and friction points this paper is trying to address.

healthcare NLP agents
large language model
clinical documentation
evidence retrieval
patient messaging
Innovation

Methods, ideas, or system contributions that make the work stand out.

workflow-aware
episode-level evaluation
healthcare NLP agents
cost-sensitive treatment
💼 Related Jobs
No related jobs found.
J
Junyi Yao
Washington University in St. Louis, USA
Baichuan Li
Baichuan Li
The Chinese University of Hong Kong
Z
Zihao Zheng
Washington University in St. Louis, USA
J
Jiayu Long
Washington University in St. Louis, USA