Alyssa Habermann, Tania Gupta, Stephen Cohen, Syed Shah, Adam Khader, Kathryn Fong, Michael F Amendola
IntroductionPatients are increasingly turning to large language models (LLMs) such as ChatGPT for medical guidance, prompting concerns around medical competence. This study assesses the proficiency of ChatGPT 4.o in answering patient questions regarding the surgical management of colon and rectal cancer.MethodsChatGPT 4.o was prompted with 10 patient-based questions regarding the surgical management of colon and rectal cancer. A five-surgeon panel evaluated each response for completeness, accuracy, and effective communication (5-point Likert scale). ANOVA was used to compare difference between response categories and unpaired t-test was used to compare scores between colon and rectal cancers.ResultsThe average scores of all colon and rectal questions were 13.02 ± 0.57 and 12.94 ± 0.55, respectively (out of 15). The highest scoring colon question related to postoperative length of stay (13.80 ± 1.30) and the lowest scoring colon question pertained to operative time (12.20 ± 1.92). The highest scoring rectal questions discussed bowel continence and return to normal activities (13.60 ± 1.34 and 13.60 ± 1.52) and the lowest scoring rectal question explained surgical procedures for removing rectal cancer (11.80 ± 2.28). Average scores (out of 5) for completeness, accuracy, and communication were 4.56 ± 0.69, 4.35 ± 0.73, and 4.16 ± 0.71, respectively (P < .001). There was no significant difference between evaluator scores for colon and rectal questions (P > .05).ConclusionsChatGPT 4.o is capable of answering commonly asked patient questions regarding colon and rectal cancer with reasonable completeness, accuracy, and effective communication with statistically significant variations between these categories.