科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Communications Medicine2026-01-09· Generalizability theory

Structured clinical approach to enable large language models to be used for improved clinical diagnosis and explainable reasoning

Muhammad Ayoub, Hai Zhao, Lifeng Li, Dongjie Yang, Shabir Hussain, Junaid Abdul Wahid

原始摘要(英文原文)· Original abstract
Applying large language models to medicine faces critical trust challenges in diagnostic reasoning. Existing approaches often fail to generalize across different models and datasets, particularly those covering a wide range of diseases and diverse patient records. This study aims to develop a universal model-based clinical framework that improves diagnostic performance while providing explainable reasoning. We introduce a structured clinical approach that replicates real-world diagnostic workflows. Patient narratives are first transformed into labeled clinical components. A validation mechanism then checks model-generated diagnoses using a disease knowledge algorithm. Additionally, a stepwise decision-making model simulates consultations progressing from junior to senior clinicians to refine diagnostic reasoning. The framework is evaluated across multiple large language models and clinical reasoning datasets using standard diagnostic accuracy metrics. Here we show that our approach outperforms existing prompting methods across six large language models and two clinical datasets. One model achieves the highest diagnostic F1 scores (0.93 on NEJM, 0.95 on MedCaseReasoning) with minimal misclassification (1 false positive and 3 false negatives). It also attains the best text-based reasoning scores on NEJM, demonstrating effective, explainable clinical outputs. When validated on real-time electronic health record data, the method shows high diagnostic accuracy (0.91) and human-like rationales (4.5 out of 5), confirming its applicability in real-world clinical settings. These findings confirm the robustness and generalizability of our framework, highlighting its potential for reliable, scalable, and explainable clinical decision support across diverse models and datasets. Medical decisions often rely on understanding complex patient information, and computational tools such as large language models can help analyze this data. However, these models sometimes make errors and their reasoning is not always clear. In this study, we developed a system that mimics how doctors work, breaking down patient notes into understandable pieces and checking model-generated diagnoses for accuracy. We tested this approach across multiple models and datasets. Here we show that it improves diagnostic accuracy, produces understandable explanations, and works well on real patient records. This method could make computer-assisted diagnosis more reliable, helping doctors make better decisions and potentially improving patient care in the future. Ayoub et al. describe a structured clinical framework that guides large language models through stepwise diagnostic reasoning, mimicking real-world clinical workflows. This approach improves diagnostic accuracy, generates human-like explanations, and generalizes across models, datasets, and real-world hospital data.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Structured clinical approach to enable large language models to be used for improved clinical diagnosis and explainable reasoning — 科研速览 Science Skim