科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in medicine2026-01-01

Reasoning vs. conventional large language models for BI-RADS educational questions answering: a multi-model comparative evaluation.

Yuxia Tang, Yuting Liu, Mengxuan Liu, Hao Ni, Siqi Wang, Shouju Wang

一句话结论 · In one sentence

Reasoning LLMs show significant potential for BI-RADS guideline explanation and education, but require specific optimization for complex clinical scenario instruction.

原始摘要(英文原文)· Original abstract
OBJECTIVE: To compare reasoning vs. conventional large language models (LLMs) in generating answers with guideline-aligned explanations for Breast Imaging Reporting and Data System (BI-RADS) educational questions. METHODS: In this prospective study performed from February 6 to 12, 2025, 49 English-Chinese question pairs were extracted from BI-RADS Atlas Fifth Edition. Two reasoning LLMs (ChatGPT-o1, Deepseek-R1) and six conventional LLMs (Gemini2.0-Flash, Deepseek-V3, ChatGPT-4o, ChatGPT-3.5, Qwen-2.5, and WenXinYiYan-3.5) generated answers and explanations to the questions through structured prompts. Three radiologists specialized in breast imaging independently evaluated responses using a 5-point Likert scale, with reference to standard answers. RESULTS: The reasoning LLMs significantly outperformed conventional models (median [interquartile range (IQR)]: 3.7 [2.7-4.0] vs. 2.7 [2.0-3.7], P < 0.001), with ChatGPT-o1 and Deepseek-R1 demonstrating peak performance. Both categories of LLMs exhibited significant score reductions in handling questions with multifaceted clinical scenarios (reasoning models: median 4.0 [2.7-4.3] vs. 2.7 [2.3-2.7], Δmedian = -1.3, P < 0.001; conventional models: 2.7 [2.0-3.7] vs. 2.3 [2.0-2.7], Δmedian = -0.4, P < 0.001). While question language showed no significant impact on reasoning LLMs (ChatGPT-o1 and Deepseek-R1), it affected some conventional models (ChatGPT-3.5, Deepseek-V3 and Gemini2.0-Flash). LLMs performance remained independent of question section and question type. CONCLUSIONS: Reasoning LLMs show significant potential for BI-RADS guideline explanation and education, but require specific optimization for complex clinical scenario instruction.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Reasoning vs. conventional large language models for BI-RADS educational questions answering: a multi-model comparative evaluation. — 科研速览 Science Skim