Noise Floor Audit for Agent Benchmarks

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过匹配AST评分审计了不同提供商的工具调用端点在官方BFCL类别中的测量变异性,揭示了语义保留提示扰动对结果稳定性的影响。
📝 Abstract
We audit measurement variability for 3 native tool-calling endpoints across 2 providers on the official BFCL multiple and parallel categories, using matched AST grading. At temperature 0, reruns are nearly deterministic across Groq endpoints and a thinking-enabled Gemini setting: ever-flip fractions are 0.7%, 2.0%, and 2.7%, with mean run correlations of 0.997, 0.966, and 0.961. Semantics-preserving prompt perturbations create the larger floor on all endpoints, with median perturbation paired SDs 11x to 58x larger than rerun paired SDs. The failure character also shifts: malformed-output failures account for 30%, 7%, and <1% of task failures, so marginal accuracy hides not only stability but also failure mode.
Problem

Research questions and friction points this paper is trying to address.

measurement variability
tool-calling endpoints
AST grading
rerun determinism
prompt perturbations
Innovation

Methods, ideas, or system contributions that make the work stand out.

measurement variability
matched AST grading
semantics-preserving prompt perturbations
🔎 Similar Papers