Sofyan Jankowski, Charbel Mourad, Marie Nowak, Jonas Richiardi, Wendy Brito Rodriguez, Florian Poncet, Marianna Gulizia, Stephanie de Labouchere, Emily Harkness, David Rotzinger, Francesco Ria, Chiara Pozzessere
Unlike experts, patients preferred ChatGPT-generated responses to institutional materials for radiology risk questions. This divergence highlights the need for patient-centered communication and suggests that large language model-based styles, implemented with expert oversight, may improve the perceived clarity and trustworthiness of educational materials in medical specialties that use ionizing radiation.
BACKGROUND: Radiology-risk communication affects multiple clinical specialties that use ionizing radiation, and many patients seek related information online. Prior expert evaluations found comparable performance between ChatGPT-generated and radiology-risk answers from official institutions, but patient perspectives have not been assessed.
PURPOSE: To assess patients' perceptions of ChatGPT versus human-generated radiology-risk information.
METHODS AND MATERIALS: From December 2024 to March 2025, patients at 3 hospitals in the United States, Switzerland, and Lebanon were randomly assigned to 1 of 5 common radiology-risk questions. Participants, blinded to source, provided subjective ratings of both ChatGPT‑3.5 and human-generated institutional responses on 7-point Likert scales for satisfaction (primary outcome), comprehensibility, trust, and reassurance. Quantitative comparisons were performed with Inverse Normalizing Transformation, and free-text comments were analyzed using thematic coding.
RESULTS: A total of 328 patients participated (34% aged 18-39 years, 33% aged 40-59 years, 31% aged 60-79 years, 3% aged ≥80 years; 188 female). ChatGPT responses were rated significantly higher than human responses for satisfaction (0.70-point advantage; P < .001), comprehensibility (0.27 points; P < .01), trust (0.65 points; P < .001), and reassurance (0.51 points; P < .001). Findings converged with qualitative written comments (r = 0.91, P < .05), in which ChatGPT attracted 2.3× more positive comments while human-generated responses received 1.7× more negative comments.
CONCLUSIONS: Unlike experts, patients preferred ChatGPT-generated responses to institutional materials for radiology risk questions. This divergence highlights the need for patient-centered communication and suggests that large language model-based styles, implemented with expert oversight, may improve the perceived clarity and trustworthiness of educational materials in medical specialties that use ionizing radiation.