科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Digital health2026-01-01

Dynamic alignment of large language models for evidence-grounded heart failure decision support.

Lu Liu, Chenchen Dong, Yunbo Ba, Haihong Yan, Xiaoxiao Tang, Yu Sun, Huilin Chen, Boyuan Shi, Qin Yu, Shulong Zhang

一句话结论 · In one sentence

Dynamic alignment shifted the model from linguistic mimicry toward clinically constrained HFrEF decision support. These findings suggest that staged optimization with policy alignment and retrieval grounding can improve evidence-based recommendations, while conventional language-overlap metrics may underestimate clinically safer generation.

原始摘要(英文原文)· Original abstract
OBJECTIVES: Large language models (LLMs) are increasingly studied for clinical decision support, but high-risk cardiology exposes persistent weaknesses in hallucination control, guideline adherence, and medication-safety reasoning. Heart failure with reduced ejection fraction (HFrEF) is a demanding test case because safe care requires structured guideline-directed therapy, comorbidity-aware monitoring, and reliable risk warnings. METHODS: We developed a dynamic alignment framework using 1087 retrospective HFrEF cases from Affiliated Zhongshan Hospital of Dalian University. An open-source LLaMA-3.1 backbone was optimized through four sequential stages: continual pre-training for heart-failure domain adaptation, supervised fine-tuning for structured clinical responses, reinforcement policy optimization for safety-oriented alignment, and retrieval-augmented generation for guideline grounding. Models were assessed with dual-track clinical and linguistic metrics. RESULTS: LLaMA-3.1 was the strongest supervised baseline, but supervised fine-tuning alone did not fully resolve guideline-adherence limitations. Staged alignment produced a measurable Alignment Tax: the final retrieval-grounded variant improved the Clinical Score from 0.716 to 0.864 and reached a Guideline Score of 0.881, while BLEU-4 decreased from 0.371 to 0.272. The decline in surface overlap coincided with stronger risk safety, stricter structure, and more guideline-directed outputs. CONCLUSIONS: Dynamic alignment shifted the model from linguistic mimicry toward clinically constrained HFrEF decision support. These findings suggest that staged optimization with policy alignment and retrieval grounding can improve evidence-based recommendations, while conventional language-overlap metrics may underestimate clinically safer generation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Dynamic alignment of large language models for evidence-grounded heart failure decision support. — 科研速览 Science Skim