API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过审计ChatGPT等模型在不同系统和基准测试中的表现,揭示了API与实际聊天界面之间存在性能差异的问题,挑战了API评分可靠性的假设。
📝 Abstract
Benchmark scores are a central currency in model releases: they inform purchasing decisions, shape public trust, and influence policy. Yet, a key assumption underlying benchmark scores is that the model performance measured through APIs faithfully reflects the behavior of deployed systems. We challenge this assumption by auditing ChatGPT, Claude, and Gemini across seven systems and nine benchmarks spanning general capability, social bias, and sycophancy. We find systematic API--interface differences in both accuracy and consistency. On average, API evaluations score 3.4 percentage points higher in accuracy and 2.1 percentage points higher in test--retest agreement than corresponding interface evaluations. For ChatGPT, the performance difference between API and interface access exceeds the API-only difference between GPT 5.3 and GPT 5.4. Put differently, switching access surfaces can degrade performance as much as downgrading a full model generation. We further test whether exposed API controls can reproduce interface behavior by varying system prompts, sampling parameters, and reasoning settings. These controls shift behavior in some cases but do not reliably eliminate the gap. Our findings document a context-validity gap: measurements obtained through APIs do not necessarily generalize to corresponding deployed interfaces, complicating the use of API evaluations as proxies for deployed systems.
Problem

Research questions and friction points this paper is trying to address.

API Benchmark Scores
Chatbot Interfaces
Model Performance
System Prompts
Context-Validity Gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

API Benchmark Scores
Interface Differences
Context-Validity Gap
🔎 Similar Papers
2024-03-12Annual Meeting of the Association for Computational LinguisticsCitations: 16