Celal Akdemir, Mücahit Furkan Balcı, Fatih Yıldırım, Mehmet Ferdi Kıncı, Ramazan Erda Pay
In this clinician-based expert review, ChatGPT® showed relatively strong performance in accuracy but lower performance in comprehensiveness and variable performance in safety when addressing FSD-related questions. These findings suggest that ChatGPT® may serve as a supplementary patient-education tool for basic, guideline-aligned topics, but its outputs should be regarded as informational starting points rather than stand-alone clinical recommendations.
OBJECTIVE: To provide a methodologically robust evaluation of ChatGPT® responses to frequently asked patient questions on female sexual dysfunction (FSD) using operationalized rating criteria and cluster-aware statistical analyses.
METHODS: This study was conducted in Türkiye. Ten commonly encountered FSD questions were identified through a structured review of patient encounters. ChatGPT® (March 2025) generated standardized responses using a fixed prompt. A multinational panel of obstetrics and gynecology specialists (n = 22) evaluated each response across three operationally defined criteria-accuracy, comprehensiveness, and safety-using a 5-point anchored Likert scale. The intentionally incorrect item was retained in the full dataset and also assessed separately as an attention check. Primary comparisons across criteria were performed using repeated-measures ANOVA, with mixed-effects sensitivity models including random intercepts for raters and questions. Secondary descriptive threshold analyses were performed using alternative positivity cut-points.
RESULTS: A total of 660 ratings were generated across ten ChatGPT® responses. The overall mean rating was 4.08 ± 0.27. Accuracy received the highest mean score (4.22 ± 0.26), followed by safety (3.96 ± 0.28) and comprehensiveness (3.86 ± 0.30). Repeated-measures ANOVA showed significant differences across the three criteria (p = 0.014), and mixed-effects modeling demonstrated the same directional pattern. Treatment-focused questions received higher ratings (4.19 ± 0.24) than diagnostic questions (3.75 ± 0.31). Inter-rater reliability was moderate to good across domains. The intentionally incorrect response was correctly identified as inaccurate and/or unsafe by 20 of 22 raters (90.9%).
CONCLUSION: In this clinician-based expert review, ChatGPT® showed relatively strong performance in accuracy but lower performance in comprehensiveness and variable performance in safety when addressing FSD-related questions. These findings suggest that ChatGPT® may serve as a supplementary patient-education tool for basic, guideline-aligned topics, but its outputs should be regarded as informational starting points rather than stand-alone clinical recommendations.