D. Austria, B. McCollister, J. E. Lindsey, M. Arowolo, M. Okon
Practitioner-derived taxonomies of conversational artificial intelligence (AI) workflow errors show low inter-rater agreement among human coders, leaving open whether the instrument is ill-specified or the judgments are inherently difficult. We delivered a locked eight-category workflow-error taxonomy verbatim, under standardized conditions, to three frontier large language model comparators from distinct developer lineages, which coded a documented 45-incident error corpus. On the same 16 incidents coded by three human reviewer-authors, comparator category agreement was substantial (Fleiss {kappa}=0.625) against slight human agreement ({kappa}=0.155); across the full corpus it was stable ({kappa}=0.632), and a 10-category refinement did not reduce it ({kappa}=0.690). Severity and a claimed-verification flag remained only fair in both arms and on both samples. Substantial cross-system consistency provides a legibility signal consistent with recoverable category distinctions, but cannot separate instrument clarity from shared model priors, and is not a validation of any coding.