Tianai Huang, Lu Lu, Jiayuan Chen, Lihao Liu, Junjun He, Yuping Zhao, Wenchao Tang, Jie Xu
Large language models (LLMs) show promise in specialized domains, but their capabilities in Traditional Chinese Medicine(TCM)—a field with complex theoretical foundations and clinical practices—remain inadequately assessed. This study aims to develop and apply a comprehensive, multidimensional benchmark to systematically evaluate LLM proficiency in this culturally-grounded medical discipline. We constructed the TCM-3CEval benchmark to assess models across three core dimensions: mastery of fundamental knowledge, comprehension of classical texts, and clinical decision-making. The dataset comprises 450 expert-validated single-choice questions. We evaluated a diverse set of models, including general-purpose international models, Chinese general models, and medical domain-specific models. Model performance was measured by accuracy under a rigorous permutation-based consistency test to control for positional bias. We show that while top-tier models approach human-level performance in aggregate scores, distinct deficits remain in clinical reliability. Models trained with Chinese linguistic and cultural priors demonstrate superior capability in interpreting classical texts compared to international counterparts. However, despite achieving comparable overall accuracy to human students (p > 0.05), models exhibit pronounced deficiencies in critical subdomains such as TCM Diagnostics and Syndrome Differentiation, where human reasoning remains significantly more robust. The benchmark effectively discriminates between rote memorization and the holistic inference required for clinical practice. This study establishes a standardized, triaxial evaluation paradigm for assessing AI in Traditional Chinese Medicine. The findings underscore the importance of cultural-contextual alignment and domain-specific adaptation for developing capable medical LLMs. This benchmark provides a foundation for optimizing models in culturally-grounded medical domains and supports the responsible integration of AI into traditional medical education and practice. Artificial intelligence (AI) language models are increasingly used in medicine. This study aimed to see how well these AI models perform in the specialized field of Traditional Chinese Medicine (TCM). We created a comprehensive test that evaluates AI on three key TCM skills: basic knowledge, understanding of ancient medical texts, and clinical decision-making. We tested a variety of AI models with this benchmark. We found that AI models trained with Chinese language and cultural knowledge performed much better at understanding TCM than general international models. However, all models still struggled with some advanced TCM topics. This research provides a clear and standard way to measure AI capabilities in TCM. The findings help guide the development of better medical AI tools, which could one day support doctors in making diagnoses and personalizing treatments, potentially improving healthcare for patients. Huang, Lu et al. develops a triaxial benchmark (TCM-3CEval) to evaluate large language models in Traditional Chinese Medicine (TCM). The findings show that models trained with Chinese cultural and linguistic knowledge outperform others, yet all exhibit critical gaps in specialized TCM subdomains.