Helia Azmakan, Tahmour Azamakan
The evaluated LLMs performed strongly in recognizing TDM problems in structured vignettes but showed variability in reasoning-intensive tasks, with overconfident reasoning representing a potential patient safety concern. Pharmacist oversight remains essential for TDM tasks.
BACKGROUND: Large language models are increasingly investigated as clinical decision support tools, but their reliability for therapeutic drug monitoring interpretation remains poorly explored. Phenytoin and digoxin, two narrow therapeutic index drugs, represent clinically challenging test cases.
OBJECTIVES: To evaluate three current-generation LLMs on phenytoin and digoxin TDM interpretation across seven clinical reasoning domains, identify domain-specific strengths and limitations, and characterize failure patterns using a qualitative error taxonomy.
METHODS: Thirty structured clinical vignettes were submitted to Claude Sonnet 4.6, ChatGPT 5.5, and Gemini 3.1 Pro. The blinded responses were independently scored by two raters using a seven-domain rubric. Between-model differences were assessed using Friedman and post-hoc Wilcoxon signed-rank tests, with Holm adjustment across domain-level tests and Bonferroni correction for pairwise comparisons. All 90 responses underwent error taxonomy analysis. Second independent responses were subsequently generated and scored using the same procedure to assess agreement across repeated generations.
RESULTS: Claude achieved the highest mean performance (94.9% ± 8.3%), significantly outperforming ChatGPT (84.5% ± 12.9%, p<0.001) and Gemini (83.4% ± 11.9%, p=0.001). All models showed high performance on level interpretation and toxicity assessment, but significant between-model differences emerged on pharmacokinetic reasoning, management, monitoring, and uncertainty acknowledgment (all Holm-adjusted p≤0.02). Overconfident reasoning was the most common error category (45.4% of total error appearances). Repeat-generation analysis showed high test-retest agreement across generations (ICC=0.90; 95% CI, 0.82-0.94).
CONCLUSION: The evaluated LLMs performed strongly in recognizing TDM problems in structured vignettes but showed variability in reasoning-intensive tasks, with overconfident reasoning representing a potential patient safety concern. Pharmacist oversight remains essential for TDM tasks.