Mert Çetin, Ali Çağlar Turgut, Orhan Beger, Fang Wang, Aaron S Dumont, Fuyou Guo, Mehmet Turgut, R Shane Tubbs
Contemporary LLMs perform variably in providing patient-oriented information about CMs. Although some AI systems generated relatively accurate and comprehensive responses for well-established CM subtypes, performance was lower for rarer and less clearly defined variants, highlighting the continuing importance of expert oversight in delivering AI-assisted medical information.
OBJECTIVE: To evaluate the performance of contemporary large language model (LLM)-based artificial intelligence (AI) systems in providing patient-oriented information across multiple Chiari malformation (CM) subtypes.
METHODS: Five AI models (ChatGPT-4o, Gemini, Copilot, DeepSeek, and Perplexity AI) were evaluated using seven patient-oriented questions on the definition, epidemiology, etiology, symptomatology, diagnosis, treatment, and prognosis of CM. Nine CM subtypes were included: Types 0, 0.5, 1, 1.5, 2, 3, 3.5, 4, and 5. A total of 315 AI-generated responses were independently assessed by three blinded neurosurgeons using a three-point ordinal scoring system for accuracy, comprehensiveness, and conciseness.
RESULTS: There were significant differences among the AI models across all domains evaluated (all p<0.001). Gemini performed most strongly in accuracy and comprehensiveness, while Perplexity performed best in conciseness. Copilot generally performed worse across the domains evaluated. Significant subtype-based differences were also identified, with AI-generated responses concerning Types 1 and 2 performing substantially better than those concerning rarer subtypes. Diagnostic questions achieved the highest overall performance, while prognosis- and definition-related questions performed comparatively less well. Reviewer-based analyses revealed substantial variability in scoring behavior.
CONCLUSIONS: Contemporary LLMs perform variably in providing patient-oriented information about CMs. Although some AI systems generated relatively accurate and comprehensive responses for well-established CM subtypes, performance was lower for rarer and less clearly defined variants, highlighting the continuing importance of expert oversight in delivering AI-assisted medical information.