科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Clinical anatomy (New York, N.Y.)2026-08-26

Generative Artificial Intelligence Performance on University-Level Human Anatomy Examinations: A Structured Narrative Review and Proposed MATRIX-Anatomy Framework.

Juan A Sanchis-Gimeno, Juan José Valenzuela-Fuenzalida, Alejandro Bruna-Mejías, Mathias Orellana-Donoso, Glen J Paton, Shahed Nalla

原始摘要(英文原文)· Original abstract
Generative artificial intelligence (GenAI) can perform strongly on written anatomy examinations, but whether such scores represent anatomical competence remains uncertain because results vary with the model, assessment, protocol, modality, scoring, and comparator. We synthesized studies evaluating GenAI as the examinee in university-level human anatomy assessments and proposed MATRIX-Anatomy, a reporting and interpretive framework not yet externally validated. PubMed, Scopus, and Web of Science Core Collection were searched for records published from January 2022 to 11 July 2026. Eligible studies used university examinations, course item banks, or curriculum-aligned undergraduate benchmarks and reported quantitative performance. Three reviewers completed study selection, data extraction, and narrative synthesis by consensus. Of 115 records, 51 duplicates were removed, 64 were screened, and 15 studies were included. Leading systems scored 76% to 98% on text-based multiple-choice assessments. On a fixed 120-item set, accuracy increased from 45.8% with ChatGPT-3.5 to 86.7% with ChatGPT-5. Human comparisons were mixed. Visuospatial performance was weaker: ChatGPT-4o identified 22.26% of cadaveric structures after up to three attempts; ChatGPT-4.0 achieved 17.3% end-to-end accuracy on image-based anatomy; and ChatGPT-5.1 reached 74.4% on a surgical-anatomy subset. Repeated runs revealed volatility and consistently incorrect responses. The findings support supervised formative use with authoritative verification and retention of secure supervised, oral, constructed-response, visuospatial, and practical assessments. MATRIX-Anatomy specifies six domains (Model, Assessment, Testing protocol, Reference standard, Input, and eXternal validity) for reproducible reporting and defensible interpretation, but requires formal external validation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Generative Artificial Intelligence Performance on University-Level Human Anatomy Examinations: A Structured Narrative Review and Proposed MATRIX-Anatomy Framework. — 科研速览 Science Skim