科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of Orthopaedic Reports2026-02-01· Readability

Evaluation of artificial intelligence-generated responses to patient inquiries for orthopaedic sports procedures

Nathan Beda, Harisankeerth Mummareddy, Ryan Qiu, Evan Porter

原始摘要(英文原文)· Original abstract
Patients are increasingly turning to online resources to guide medical decision-making, particularly younger populations at risk for orthopedic injuries (2). Large language models (LLMs) such as ChatGPT and Google Gemini have emerged as alternatives to traditional search engines, offering synthesized responses rather than static web page links. The readability, length, and source transparency of LLM-generated responses in the field Orthopaedic Surgery continue to be explored in comparison to conventional Google Search. To evaluate and compare the readability, response length, and source characteristics of patient-focused orthopedic information generated by ChatGPT, Google Gemini, and Google Search. Cross-sectional, comparative analysis of online information sources. A total of 40 standardized orthopedic-related patient questions were posed to ChatGPT 4.0, Google Gemini, and Google Search. Responses were assessed for word count, Flesch Reading Ease (FRE), and Flesch–Kincaid Grade Level (FKGL). Google Search sources were categorized using the Modified Rothwell Criteria, which classifies websites as commercial, academic, medical practice, single surgeon, government, or social media. Pairwise t-tests were performed to compare readability and response length across platforms, with statistical significance set at p < 0.05. LLM responses were significantly longer and more complex than Google Search snippets. Average response lengths were 342.75 words for ChatGPT 4.0, 306.88 words for Google Gemini, and 40.18 words for Google Search (p < 0.001). FRE scores indicated difficult readability for ChatGPT (25.6) and Gemini (26.0) versus a significantly easier comprehensibility for Google Search (40.7, p < 0.05). FKGL analysis showed ChatGPT responses required a higher reading level (13.7) than Google Search (12.6, p < 0.05). Source analysis of Google Search revealed that 55% of results were from academic sites, 32.5% from medical practices, 7.5% from single surgeons, 2.5% from government websites, and 2.5% from social media, with no commercial websites represented. LLMs did not provide explicit source citations. Patients increasingly rely on LLMs for orthopedic information, specifically the younger generation most at risk for sports injury (5). While LLM responses provide greater detail and context than traditional search results, their higher reading complexity and lack of transparent sourcing may challenge comprehension and limit the effectiveness of patient education. Clinicians should consider these factors when integrating AI tools into shared decision-making and patient counseling.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluation of artificial intelligence-generated responses to patient inquiries for orthopaedic sports procedures — 科研速览 Science Skim