Kemal Turker Ulutaş, Adnan Pekmezci
Background/Objective: Large language models (LLMs) have shown exam-level performance, yet their reliability and safety in laboratory medicine-where quantitative data interpretation is central-remain insufficiently validated. This study compared the accuracy, interpretive quality, and safety of ChatGPT-5.2, Gemini 3 Pro, and DeepSeek-V3.2 using a standardized, text-based educational benchmark of clinical pathology and laboratory medicine vignettes. Methods: For each case, the original open-ended questions were answered by each model and scored by two blinded expert raters across four domains-diagnostic accuracy, interpretation, management/investigations, and safety-using a six-point Likert scale, yielding a composite score ranging from 4 to 24. Investigators developed five single-best-answer MCQs per case (500 MCQs total) with consensus answer keys; models selected one option per item. Results: Inter-rater agreement was high (κ = 0.80 for diagnostic concordance; κ = 0.78 for safety flags). Mean composite open-ended scores were 22.5 ± 2 for ChatGPT-5.2, 21.1 ± 2.5 for Gemini, and 20.6 ± 2.6 for DeepSeek (p < 0.001). Fully concordant primary diagnoses were observed in 88%, 84% and 82% of cases, respectively. Unsafe or guideline-discordant recommendations were uncommon but present (2%, 4%, 5%). MCQ accuracy was high: 96% (480/500) for ChatGPT-5.2, 95% (475/500) for Gemini, and 94% (470/500) for DeepSeek. Conclusions: All three LLMs achieved high performance in this standardized text-based benchmark. ChatGPT-5.2 had the highest observed performance across several predefined outcomes, although absolute differences were modest for some measures, particularly MCQ accuracy. These findings should not be interpreted as evidence of clinical superiority or real-world effectiveness.