Azfar Athar Ishaqui, Tauqeer Hussain Malhi, Emad Ali Alsaleh, Rayah Asiri, Saira Faraz Shah, Amer Hayat Khan, Khalid Orayj, Salman Ashfaq Ahmad, Muhammad Imran, Narendar Kumar, Adnan Iqbal, Muhammad Bilal Maqsood
These findings represent comparative model benchmarking and require validation against pharmacist-led therapeutic drug monitoring and Bayesian AUC-guided dosing platforms before clinical use.
BACKGROUND AND OBJECTIVES: Vancomycin monitoring remains challenging due to fluctuating renal function and ICU physiology. This study compares ChatGPT and Grok for predicting sequential vancomycin trough levels and dosing intervals across varying renal function.
RESEARCH DESIGN AND METHODS: This retrospective study used deidentified MIMIC-IV ICU data. Admissions with three sequential vancomycin troughs and complete dosing history were included (239 admissions, 717 predictions: T1-T3). Structured clinical snapshots were submitted to both models to predict trough concentrations (mg/L) and dosing intervals (hours). Performance was assessed using MAE, RMSE, bias, ±2 mg/L and ±2-hour accuracy, sequence success, and generalized estimating equations.
RESULTS: ChatGPT outperformed Grok in trough prediction (MAE 4.07 vs 5.08 mg/L; RMSE 5.55 vs 7.53 mg/L; ±2 mg/L accuracy 37.0% vs 29.0%), with greater advantage at T1/T2 and in renal dysfunction. Grok was superior for interval prediction (MAE 1.36 vs 1.61 hours; RMSE 2.09 vs 2.98 hours). ChatGPT achieved more successful trough sequences (≥2/3 within ±2 mg/L: 35.1% vs 23.0%), while Grok had more accurate interval sequences (all 3 within ±2 hours: 67.8% vs 59.8%).
CONCLUSIONS: These findings represent comparative model benchmarking and require validation against pharmacist-led therapeutic drug monitoring and Bayesian AUC-guided dosing platforms before clinical use.