Wenfu Zhong, Jingwen Huang, Mengsi Wei, Qingpeng Liang, Yuanwu Yang, Jinrong Chen, Jinjiang Mao, Ying Qin, Yifan Sun, Yishan Liang
In regions with high thalassemia prevalence, such as southern China and Southeast Asia, chronic shortages of professional genetic counseling resources have driven interest in large language models (LLMs) as auxiliary tools, yet their performance and safety boundaries in this setting remain uncharacterized. This single-center retrospective study evaluated four LLMs (ChatGPT-5.2 Thinking, DeepSeek-V3.2 chat, Gemini 3 Flash, and Grok 4.1 Fast) using 1,080 standardized knowledge questions administered across five independent sessions and 150 real-world clinical cases scored by six senior experts across eight dimensions. All models exceeded 90% accuracy on single-choice and true-false questions. Between-model differences were most pronounced in multiple-choice questions, where ChatGPT-5.2 Thinking achieved the highest accuracy (87.28% ± 1.69%), significantly outperforming Grok 4.1 Fast (72.06% ± 1.69%, P < 0.001). DeepSeek-V3.2 chat showed lower cross-session consistency than the other three models. In clinical case analysis, all models scored below human expert levels overall, with ChatGPT-5.2 Thinking performing closest to experts. Test report interpretation was generally adequate, whereas larger gaps emerged in genetic risk estimation, phenotype prediction, and counseling recommendations (all P < 0.001 vs. human experts, with limited exceptions for ChatGPT-5.2 Thinking in β-thalassemia subtypes). All 145 severe errors were confined to α-thalassemia and α-combined-β-thalassemia subtypes, concentrated in phenotype prediction (76/145, 52.4%) and risk estimation (40/145, 27.6%). LLMs may support knowledge retrieval and structured report interpretation in thalassemia genetic counseling but should not be used independently for risk assessment or final counseling decisions, particularly in complex α-related cases.