Fatemeh Shams, Elina Khanehzar, Amirsajad Jafari, Alireza Poustforoosh, Masoud Abedi, Sonia Jafarinia
This study introduces a novel framework, the SANDI checklist (Synthetic Accessibility, ADME, Novelty, Docking Score, Improvement), to systematically evaluate the performance of large language models (LLMs) in computational drug design. Four LLMs, DeepSeek, ChatGPT, Gemini, and Grok, were tasked with designing 20 de novo small-molecule inhibitors of the AXL receptor tyrosine kinase (n = 20 molecules per model), guided by optimized prompts that enforced scaffold diversity, drug-likeness, and favorable pharmacokinetic profiles. All models were queried using a fixed prompt in single-generation runs in the web GUI, and all findings are based exclusively on in-silico analyses, with predicted toxicity flags reported as computational risk indicators. Molecular docking and dynamics simulations assessed binding affinities and interaction stability with AXL (PDB ID: 5U6B ), using FDA-approved drugs (e.g., bemcentinib, cabozantinib) as controls. All results reflect single-generation runs and in silico predictions only. DeepSeek achieved the highest overall SANDI score (80.43%), excelling in docking (79.8%) and synthetic accessibility (81.5%), while Grok led in novelty (85%). Refinements of top compounds generated using the DeepSeek model with each of the four LLMs resulted in the GPTR1 compound, which exhibited superior interaction profiles in molecular dynamics, forming the most hydrogen bonds and effectively stabilizing AXL residues. These findings demonstrate the utility of SANDI as a benchmarking framework rather than evidence of clinical efficacy. The SANDI framework offers a standardized approach to evaluate LLMs in drug discovery, revealing their strengths and limitations in generating viable inhibitors and paving the way for refined AI-driven drug design strategies.