科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ arXiv2026-09-08· cs.AI

CIVI: A Framework for Diagnosing Search Agent Failures in Civic Information

Dingying Liu, Yunshun Zhong, Wentao Zhang, Yiyuan Li

原始摘要(英文原文)· Original abstract
Large Language Models are increasingly deployed in public-sector settings, where incorrect guidance can cause irreversible harm. We introduce CIVI, the first framework for diagnosing search agent failures in civic information. Its benchmark instantiation jointly spans cross-national, interjurisdictional government contexts (federal, state, and local) and functional categories from an internationally adopted United Nations standard. We evaluate ten frontier search agents and find that none matches an attentive human baseline. Alongside accuracy, CIVI measures search invocation rate, selective no-search accuracy, and how often agents cite authoritative government sources. To perform this diagnosis, we introduce ARISE, which decomposes agentic search failures into four mutually exclusive modes, isolated via source-injection ablation. ARISE attributes 72.1% of all observed failures to retrieval-bound causes rather than to gaps in the models' parametric knowledge.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

CIVI: A Framework for Diagnosing Search Agent Failures in Civic Information — 科研速览 Science Skim