科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of the American Medical Informatics Association : JAMIA2026-09-17

DiagnosticXchange: an open-source framework for evaluating safety, efficiency, and diagnostic reasoning in clinical AI systems.

Moran Sorka, Alon Gorenshtein, Hillel Abramovitch, Pannathat Soontrapa, Shahar Shelly, Dvir Aran

一句话结论 · In one sentence

Accuracy-only evaluation is insufficient for safe clinical AI deployment. DiagnosticXchange provides an open-source, reproducible framework for multi-dimensional assessment that reveals clinically consequential differences invisible to existing benchmarks.

原始摘要(英文原文)· Original abstract
OBJECTIVE: To develop and validate an open-source evaluation framework that assesses clinical AI diagnostic systems across multiple clinically relevant dimensions (accuracy, cost, time, invasiveness, physician effort, and safety behaviors), addressing the limitations of accuracy-only benchmarks. MATERIALS AND METHODS: We developed DiagnosticXchange, a dynamic clinical simulation platform where AI systems interact with a simulated hospital by ordering tests, requesting imaging, and performing procedures, with each action mapped to Current Procedural Terminology (CPT) codes capturing cost, time, work relative value units, and invasiveness. We validated the framework using 8 large language models on 216 peer-reviewed cases across 19 specialties (1728 sessions). Analyses included competing-risk survival analysis, unsupervised clustering of reasoning strategies, diagnostic contribution scoring, and a pilot comparison with 14 neurologists. RESULTS: Three systems achieved near-identical accuracy (93.5%-94.0%) yet differed significantly in cost (P <.001; 1.75-fold between the most and least expensive top-3 systems) and 2.1-fold in physician oversight. Competing-risk analysis showed the most efficient system solved 86% of cases within $5000, while the most resource-intensive required $10 000 for equivalent accuracy. Unsupervised clustering identified 3 reasoning strategies; behavioral profiles predicted resource consumption better than system identity (AIC: 4683 vs 5056). Safety analysis revealed premature diagnosis (up to 9.3%), noncontributory invasive procedures (8.7%-29.9%), and futile invasive procedures on failed cases (3.7%-18.1%). Test-retest analysis confirmed reproducible accuracy (84.6% concordance) but substantial process variability (cost CV: 86.6%). A pilot comparison captured multi-dimensional performance of 14 neurologists alongside AI systems. CONCLUSIONS: Accuracy-only evaluation is insufficient for safe clinical AI deployment. DiagnosticXchange provides an open-source, reproducible framework for multi-dimensional assessment that reveals clinically consequential differences invisible to existing benchmarks.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

DiagnosticXchange: an open-source framework for evaluating safety, efficiency, and diagnostic reasoning in clinical AI systems. — 科研速览 Science Skim