科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ arXiv2026-08-23· cs.CL

Noise Floor Audit for Agent Benchmarks

Yihang Chen, Pin Qian, Su Wang, Chong Peng, Huan Xu, Xiyang Wu, Yiqi Sun

原始摘要(英文原文)· Original abstract
We audit measurement variability for 3 native tool-calling endpoints across 2 providers on the official BFCL multiple and parallel categories, using matched AST grading. At temperature 0, reruns are nearly deterministic across Groq endpoints and a thinking-enabled Gemini setting: ever-flip fractions are 0.7%, 2.0%, and 2.7%, with mean run correlations of 0.997, 0.966, and 0.961. Semantics-preserving prompt perturbations create the larger floor on all endpoints, with median perturbation paired SDs 11x to 58x larger than rerun paired SDs. The failure character also shifts: malformed-output failures account for 30%, 7%, and <1% of task failures, so marginal accuracy hides not only stability but also failure mode.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Noise Floor Audit for Agent Benchmarks — 科研速览 Science Skim