科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Phlebology2026-09-23

Evaluation of the accuracy and reproducibility of large language models (ChatGPT, DeepSeek, Gemini) in responding to patient-centered lipedema questions.

Rabia Sanır, Esra Nur Türkmen, Esra Giray, Figen Ayhan, Gülseren Akyüz

原始摘要(英文原文)· Original abstract
BackgroundLipedema is a frequently misdiagnosed chronic condition that significantly impacts patients' quality of life. As artificial intelligence (AI)-based large language models (LLMs) become increasingly integrated into healthcare communication, their accuracy and consistency in providing patient-centered information require thorough evaluation, especially in rare diseases like lipedema. Therefore, this study aimed to evaluate the accuracy and reproducibility of responses generated by ChatGPT, DeepSeek, and Gemini to questions frequently asked by patients with lipedema.MethodsThis cross-sectional study assessed the accuracy and reproducibility of responses generated by ChatGPT, DeepSeek, and Gemini to 25 commonly asked lipedema-related questions. Each model was queried twice in separate sessions, and answers were evaluated by three independent experts using a four-point rating scale. To ensure the objectivity and consistency of expert evaluations, inter-rater agreement was assessed using Cohen's kappa coefficient.ResultsDeepSeek achieved the highest proportion of comprehensive and correct responses (72%), followed by Gemini (64%) and ChatGPT (56%). Accuracy varied across content categories, with notable limitations particularly in treatment, follow-up, and maintenance questions. Reproducibility analysis revealed that DeepSeek produced the most consistent responses across sessions, while ChatGPT and Gemini showed more variability, particularly in treatment and quality-of-life questions. Cohen's kappa values indicated high inter-rater agreement overall, with perfect agreement in some categories for ChatGPT and DeepSeek.ConclusionsLLMs can provide generally accurate and consistent responses to patient-centered questions about lipedema, particularly in areas related to general information and diagnosis. However, reduced accuracy and reproducibility in complex clinical domains suggest that expert oversight is essential when using these tools for patient education.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluation of the accuracy and reproducibility of large language models (ChatGPT, DeepSeek, Gemini) in responding to patient-centered lipedema questions. — 科研速览 Science Skim