科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Diagnostics (Basel, Switzerland)2026-09-18

Laboratory Medicine Decision Support-Beyond Exam Passing: A Blinded 100-Case Text-Based Benchmark of Diagnostic Accuracy, Management Quality, and Safety for ChatGPT, Gemini, and DeepSeek-LLM Decision Support in Laboratory Medicine.

Kemal Turker Ulutaş, Adnan Pekmezci

原始摘要(英文原文)· Original abstract
Background/Objective: Large language models (LLMs) have shown exam-level performance, yet their reliability and safety in laboratory medicine-where quantitative data interpretation is central-remain insufficiently validated. This study compared the accuracy, interpretive quality, and safety of ChatGPT-5.2, Gemini 3 Pro, and DeepSeek-V3.2 using a standardized, text-based educational benchmark of clinical pathology and laboratory medicine vignettes. Methods: For each case, the original open-ended questions were answered by each model and scored by two blinded expert raters across four domains-diagnostic accuracy, interpretation, management/investigations, and safety-using a six-point Likert scale, yielding a composite score ranging from 4 to 24. Investigators developed five single-best-answer MCQs per case (500 MCQs total) with consensus answer keys; models selected one option per item. Results: Inter-rater agreement was high (κ = 0.80 for diagnostic concordance; κ = 0.78 for safety flags). Mean composite open-ended scores were 22.5 ± 2 for ChatGPT-5.2, 21.1 ± 2.5 for Gemini, and 20.6 ± 2.6 for DeepSeek (p < 0.001). Fully concordant primary diagnoses were observed in 88%, 84% and 82% of cases, respectively. Unsafe or guideline-discordant recommendations were uncommon but present (2%, 4%, 5%). MCQ accuracy was high: 96% (480/500) for ChatGPT-5.2, 95% (475/500) for Gemini, and 94% (470/500) for DeepSeek. Conclusions: All three LLMs achieved high performance in this standardized text-based benchmark. ChatGPT-5.2 had the highest observed performance across several predefined outcomes, although absolute differences were modest for some measures, particularly MCQ accuracy. These findings should not be interpreted as evidence of clinical superiority or real-world effectiveness.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Laboratory Medicine Decision Support-Beyond Exam Passing: A Blinded 100-Case Text-Based Benchmark of Diagnostic Accuracy, Management Quality, and Safety for ChatGPT, Gemini, and DeepSeek-LLM Decision Support in Laboratory Medicine. — 科研速览 Science Skim