科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in digital health2026-01-01· Computer science

Performance and limitations of four large language models in genetic counseling for thalassemia.

Wenfu Zhong, Jingwen Huang, Mengsi Wei, Qingpeng Liang, Yuanwu Yang, Jinrong Chen, Jinjiang Mao, Ying Qin, Yifan Sun, Yishan Liang

原始摘要(英文原文)· Original abstract
In regions with high thalassemia prevalence, such as southern China and Southeast Asia, chronic shortages of professional genetic counseling resources have driven interest in large language models (LLMs) as auxiliary tools, yet their performance and safety boundaries in this setting remain uncharacterized. This single-center retrospective study evaluated four LLMs (ChatGPT-5.2 Thinking, DeepSeek-V3.2 chat, Gemini 3 Flash, and Grok 4.1 Fast) using 1,080 standardized knowledge questions administered across five independent sessions and 150 real-world clinical cases scored by six senior experts across eight dimensions. All models exceeded 90% accuracy on single-choice and true-false questions. Between-model differences were most pronounced in multiple-choice questions, where ChatGPT-5.2 Thinking achieved the highest accuracy (87.28% ± 1.69%), significantly outperforming Grok 4.1 Fast (72.06% ± 1.69%, P < 0.001). DeepSeek-V3.2 chat showed lower cross-session consistency than the other three models. In clinical case analysis, all models scored below human expert levels overall, with ChatGPT-5.2 Thinking performing closest to experts. Test report interpretation was generally adequate, whereas larger gaps emerged in genetic risk estimation, phenotype prediction, and counseling recommendations (all P < 0.001 vs. human experts, with limited exceptions for ChatGPT-5.2 Thinking in β-thalassemia subtypes). All 145 severe errors were confined to α-thalassemia and α-combined-β-thalassemia subtypes, concentrated in phenotype prediction (76/145, 52.4%) and risk estimation (40/145, 27.6%). LLMs may support knowledge retrieval and structured report interpretation in thalassemia genetic counseling but should not be used independently for risk assessment or final counseling decisions, particularly in complex α-related cases.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Performance and limitations of four large language models in genetic counseling for thalassemia. — 科研速览 Science Skim