Longguo Zhang, Shijie Fan, Yuvathi Arunkumar, Sehrish Noor, Zhaodong Bi
AI chatbots generate accurate, complete, and consistent PFPS information but uniformly fail readability benchmarks. Disclaimer rates remain low, particularly for Gemini and Perplexity. These findings suggest AI chatbots currently function better as supplementary educational tools, highlighting the need for linguistic simplification and improved safety signaling.
BACKGROUND: Patellofemoral Pain Syndrome (PFPS) is a highly prevalent musculoskeletal condition affecting young adults and athletes. Patients increasingly turn to AI chatbots for medical information, yet the reliability, safety, and readability of these tools for PFPS remain unclear.
OBJECTIVE: To evaluate accuracy, clarity, completeness, consistency, readability, and health advice disclaimers in responses from four AI chatbots (ChatGPT, Gemini, Claude, Perplexity).
METHODS: On February 18, 2026, thirty common PFPS questions were submitted to four AI models. Anonymized responses were independently evaluated for: information quality (accuracy, clarity, completeness, consistency) using a 4-point Likert scale; readability via seven indices benchmarked against the sixth-grade level; and safety signaling by dichotomous coding of health advice disclaimers.
RESULTS: All models achieved median scores of 4.00 for completeness (P = 0.296) and consistency (P = 0.1). Significant differences emerged in accuracy (P = 0.019) and clarity (P < 0.001) overall, though adjusted pairwise accuracy differences were non-significant. Perplexity (median 3.00) was significantly inferior in clarity compared to other models (median 4.00). No model met the sixth-grade readability benchmark (P < 0.001); Gemini and ChatGPT were most readable, while Claude and Perplexity produced the most complex text. Health advice disclaimers appeared in 46.7% of ChatGPT and 40.0% of Claude responses, but only 16.7% for Gemini and Perplexity.
CONCLUSIONS: AI chatbots generate accurate, complete, and consistent PFPS information but uniformly fail readability benchmarks. Disclaimer rates remain low, particularly for Gemini and Perplexity. These findings suggest AI chatbots currently function better as supplementary educational tools, highlighting the need for linguistic simplification and improved safety signaling.