Qian Li, Yongxin Li, Chao Ye, Kaihua Zhang, Jing Bai, Xinbo Liu, Zhantao Li, Xuanyu Meng, Xianjun Min, Jin Guo
Embedding domain-specific knowledge into LLMs may improve performance on specialized exam-style thoracic-surgery questions on this text-only benchmark. However, the present 56-item evaluation does not establish clinical equivalence, diagnostic accuracy in practice, multimodal competence, or readiness for real-world clinical decision support.
BACKGROUND: Specialized thoracic-surgery questions require the integration of multi-factor clinical relationships within text, yet general-purpose large language models (LLMs) may underperform on such exam-style benchmarks.
METHODS: We constructed a DK-LLM agent by embedding curated medical textbook knowledge into a LangChain-based framework to support domain-specific reasoning. The model was tested on a 56-item thoracic-surgery examination question set in a restricted text-only setting without internet browsing or external tools and was compared with generic LLM baselines, three thoracic surgeons, and three non-expert engineers. Examination score and error patterns were assessed.
RESULTS: The knowledge-augmented DK-LLM configuration showed an 11.8-point examination-score advantage over the version without the local knowledge base. Commercial LLM-based agents outperformed the open-source baselines and non-expert participants on this question set, whereas experienced thoracic surgeons achieved the highest scores overall; in the ablation analysis, removing the local knowledge base reduced the examination score by 11.8 percentage points.
CONCLUSION: Embedding domain-specific knowledge into LLMs may improve performance on specialized exam-style thoracic-surgery questions on this text-only benchmark. However, the present 56-item evaluation does not establish clinical equivalence, diagnostic accuracy in practice, multimodal competence, or readiness for real-world clinical decision support.