科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ PloS one2026-01-01

Comparative evaluation of ChatGPT and gemini responses to patient-oriented questions on breast cancer.

Yeliz Yilmaz Bozok, Nihan Acar, Murat Kemal Atahan, Cem Karaali

一句话结论 · In one sentence

Both models demonstrated acceptable overall performance. However, ChatGPT's higher scores and shorter, more focused responses indicate that it may be a more efficient tool for addressing breast cancer-related patient questions. In their current form, these models should not be used for diagnostic or clinical decision-making purposes. AI-generated health information should be used only under expert supervision and for educational or supportive purposes.

原始摘要(英文原文)· Original abstract
BACKGROUND: Breast cancer, the most common malignancy among women, remains a major global public health concern. With the rapid growth of artificial intelligence-based language models, it is essential to evaluate their potential roles in patient education. This study compared ChatGPT and Gemini in responding to patient-oriented questions about breast cancer and its surgical treatment regarding scientific accuracy, clarity, and unnecessary detail. METHODS: Forty frequently asked questions were collected from Turkish online sources. Both models were queried under identical conditions on August 13, 2025, and the responses were anonymized for blinded evaluation. Four general surgeons experienced in breast surgery independently assessed each response using a five-point Likert scale across three domains: scientific accuracy, clarity, and unnecessary detail. RESULTS: Response lengths were recorded and compared. ChatGPT achieved significantly higher median scores than Gemini in scientific accuracy [4.75 (3.50-5.00) vs. 4.25 (3.50-5.00); p < 0.001], clarity [4.75 (3.50-5.00) vs. 4.25 (3.25-5.00); p = 0.005], and unnecessary detail [5.00 (5.00-5.00) vs. 4.50 (3.50-5.00); p < 0.001]. The overall median score was 4.83 (4.08-5.00) for ChatGPT and 4.33 (3.50-4.92) for Gemini (p < 0.001). Gemini's responses were significantly longer (244 ± 84 vs. 170 ± 47 words; p < 0.001). CONCLUSION: Both models demonstrated acceptable overall performance. However, ChatGPT's higher scores and shorter, more focused responses indicate that it may be a more efficient tool for addressing breast cancer-related patient questions. In their current form, these models should not be used for diagnostic or clinical decision-making purposes. AI-generated health information should be used only under expert supervision and for educational or supportive purposes.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Comparative evaluation of ChatGPT and gemini responses to patient-oriented questions on breast cancer. — 科研速览 Science Skim