科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ medRxiv2026-09-07· health informatics

Cross-System Legibility of a Practitioner-Derived Workflow-Error Taxonomy for Conversational AI: A Three-Comparator Agreement Study

D. Austria, B. McCollister, J. E. Lindsey, M. Arowolo, M. Okon

原始摘要(英文原文)· Original abstract
Practitioner-derived taxonomies of conversational artificial intelligence (AI) workflow errors show low inter-rater agreement among human coders, leaving open whether the instrument is ill-specified or the judgments are inherently difficult. We delivered a locked eight-category workflow-error taxonomy verbatim, under standardized conditions, to three frontier large language model comparators from distinct developer lineages, which coded a documented 45-incident error corpus. On the same 16 incidents coded by three human reviewer-authors, comparator category agreement was substantial (Fleiss {kappa}=0.625) against slight human agreement ({kappa}=0.155); across the full corpus it was stable ({kappa}=0.632), and a 10-category refinement did not reduce it ({kappa}=0.690). Severity and a claimed-verification flag remained only fair in both arms and on both samples. Substantial cross-system consistency provides a legibility signal consistent with recoverable category distinctions, but cannot separate instrument clarity from shared model priors, and is not a validation of any coding.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Cross-System Legibility of a Practitioner-Derived Workflow-Error Taxonomy for Conversational AI: A Three-Comparator Agreement Study — 科研速览 Science Skim