Natansh D. Modi, Cyril A. Alex, Abdulhalim A. Awaty, Bradley D. Menz, S. Bacchi, Kacper Gradoń, Jessica M. Logan, Andrew Rowland, Lisa M. Kalisch Ellett, Ross A. McKinnon, Michael D. Wiese, Michael J. Sorich, Ashley M. Hopkins
Unlabelled: This cross-sectional evaluation of six consumer-facing large language model platforms found significant heterogeneity in safeguard performance against the generation of health disinformation, with Claude and ChatGPT demonstrating complete resistance across all prompt types, while Copilot, Meta AI, Grok, and Gemini exhibited substantial vulnerabilities.