科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ npj Digital Medicine2026-04-15· Computer science

The effects of multitype prompt engineering for large language models in hypertension treatment decisions

Zeyan Li, Henyang Liu, Wuping Tan, Du Tang, Shoupeng Duan, Bowen Zhou, Long Tang, Xuyang Hu, Liying Huang, Peng Zhao, Wenqiang Fang, Bing Wu, Jinjun Liu, Yijun Wang, Jun Wang

原始摘要(英文原文)· Original abstract
The effects of various prompt engineering on Large Language Models (LLMs) performance in hypertension decision-making are not yet fully understood. We evaluate the impact of different prompt engineering on LLM performance in hypertension treatment decision-making. We conducted a two-stage validation study using 300 de-identified simulated hypertension cases based on real-world clinical scenarios. ChatGPT-4.1 with Guidance-Self-Consistency achieved optimal performance (91.3% accuracy), nearing expert-level competency, while zero-shot prompting yielded worst results (62.7% with DeepSeek-V3). Optimal LLM assistance consistently enhanced physicians' average accuracy across all levels (community hospital: 73.4% to 82.5%; county hospital: 84.0% to 87.9%; teaching hospital: 91.5% to 92.0%) and reduced inappropriate regimen rates. The worst LLM configurations decreased physician performance below baseline, increasing inappropriate regimen rates from 26.6% to 35.2% across all levels. Effectively designed prompt strategies enable LLMs to provide reliable hypertension treatment recommendations, thereby supporting physicians' clinical decisions. This study has been trial-registered (ChiCTR2500099307, March 21, 2025).
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

The effects of multitype prompt engineering for large language models in hypertension treatment decisions — 科研速览 Science Skim