Challenges of Auditing: Variability in Outputs of Large Language Models for Health

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究发现不同访问模式下大型语言模型在健康建议上的输出差异,强调了模型提供者需确保消费者体验和设置的忠实复制以进行严格审计。
📝 Abstract
People increasingly use frontier AI models for health advice, but via different access modes (e.g., ChatGPT, ChatGPT Health, APIs) with varying settings. Here, we find systematic differences across access modes. Because evaluations typically rely on APIs while consumers interact through chatbot interfaces, these discrepancies limit evaluation validity. Our findings underscore an urgent need for model providers to enable faithful replication of consumer experiences and settings for rigorous audits.
Problem

Research questions and friction points this paper is trying to address.

Auditing
Large Language Models
Health Advice
Access Modes
Evaluation Validity
Innovation

Methods, ideas, or system contributions that make the work stand out.

systematic differences
access modes
evaluation validity
rigorous audits
🔎 Similar Papers
No similar papers found.
Y
Yuan Pu
Department of Computer Science, Duke University, 308 Research Drive, Durham, 27708, NC.
Y
Yewon Chang
Department of Computer Science, Duke University, 308 Research Drive, Durham, 27708, NC.
F
Furong Jia
Department of Computer Science, Duke University, 308 Research Drive, Durham, 27708, NC.
Xunjian Yin
Xunjian Yin
Peking University
LLMAgentReasoning
J
Jessica Ma
Department of Medicine, Duke University, 40 Duke Medicine Circle, Durham, 27710, NC. and Geriatrics and Extended Care, Durham VA Health System, 508 Fulton Street, Durham, 27705, NC.
A
Ayman Ali
Department of Surgery, Duke University, 2301 Erwin Road, Durham, 27710, NC.
Monica Agrawal
Monica Agrawal
Assistant Professor, Duke