科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ American journal of health-system pharmacy : AJHP : official journal of the American Society of Health-System Pharmacists2026-09-08

Simultaneous ChatGPT-4o outputs for type 2 diabetes pharmacotherapy: Accuracy, usefulness, and impact variability.

Joshua Caballero, Beth Bryles Phillips, Rebecca H Stone, Robin Southwood, Sharmon P Osae, Devin Lavender, Michelle McElhannon, Katie Smith, Richard Lamb, Lorenzo Villa Zapata, Russ Palmer

一句话结论 · In one sentence

While outputs were generally accurate and useful, limitations and inconsistencies were noted. Users should be aware simultaneously generated outputs across multiple devices using ChatGPT can vary in their accuracy and usefulness in diabetes management.

原始摘要(英文原文)· Original abstract
PURPOSE: The primary objective of the study was to determine how clinically accurate and useful are ChatGPT-4o-generated outputs when identical prompts are entered simultaneously in three independent ChatGPT sessions utilizing the same user-facing model. The secondary objective was to identify if differences in outputs affect clinical outcomes (i.e., impact variability). METHODS: Five clinical prompts were developed focusing on diabetes management and counseling. Each clinical prompt was simultaneously inputted into three separate devices using ChatGPT-4o to generate outputs. A modified Delphi technique was then utilized involving five diabetes management clinical pharmacists. Each clinical pharmacy faculty independently rated each ChatGPT-generated output based on accuracy (i.e., poor, borderline, good) usefulness (i.e., not useful, somewhat useful, very useful) and impact variability (i.e, low, moderate, high). After initial assessment, responses were collated and anonymously shared among the pharmacy faculty. The faculty members were invited to revise their evaluations based on collective feedback. Pharmacy faculty then convened in a virtual panel with moderators to discuss evaluations and work towards consensus. RESULTS: Consensus was achieved for all ChatGPT outputs. Accuracy ratings ranged from borderline to good. Two of the clinical prompts yielded outputs receiving different accuracy ratings which may have impacted variability. Usefulness ratings ranged from somewhat useful to very useful, with one clinical prompt yielding outputs that received different usefulness ratings. CONCLUSION: While outputs were generally accurate and useful, limitations and inconsistencies were noted. Users should be aware simultaneously generated outputs across multiple devices using ChatGPT can vary in their accuracy and usefulness in diabetes management.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Simultaneous ChatGPT-4o outputs for type 2 diabetes pharmacotherapy: Accuracy, usefulness, and impact variability. — 科研速览 Science Skim