科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Pharmaceutical research2026-09-11

Employing General-Purpose and Biomedical Large Language Models with Advanced Prompt Engineering for Pharmacoepidemiologic Study Design.

Xinyao Zhang, Nicole Sonne Heckmann, Manuela Del Castillo Suero, Francesco Paolo Speca, Maurizio Sessa

一句话结论 · In one sentence

Off-the-shelf general-purpose LLMs currently offered more reliable support for pharmacoepidemiologic study design than the smaller biomedical LLMs evaluated. In the model-controlled comparison (GPT-4o with Least-to-Most versus Active prompting), prompt strategy did not significantly affect performance (paired p = 0.93), and prompt comparisons are therefore reported descriptively. Because the two model groups also differed in scale and instruction-tuning, this contrast should be interpreted as a comparison of readily deployable options rather than of biomedical specialization alone.

原始摘要(英文原文)· Original abstract
BACKGROUND: The potential of large language models (LLMs) to automate and support pharmacoepidemiologic study design is an emerging area of interest, yet their reliability remains insufficiently characterized. General-purpose LLMs often display inaccuracies, while the comparative performance of specialized biomedical LLMs in this domain remains unknown. METHODS: This study evaluated general-purpose LLMs (GPT-4o and DeepSeek-R1) versus biomedically fine-tuned LLMs (QuantFactory/Bio-Medical-Llama-3-8B-GGUF and Irathernotsay/qwen2-1.5B-medical_qa-Finetune) using 46 protocols (2018-2024) from the HMA-EMA Catalogue and Sentinel System. Performance was assessed across relevance, logic of justification, and ontology-code agreement across multiple coding systems using Least-to-Most (LTM) and Active Prompting strategies. RESULTS: GPT-4o and DeepSeek-R1 paired with LTM prompting achieved the highest relevance and logic of justification scores, with GPT-4o-LTM reaching a median relevance score of 4 in 8 of 9 questions for HMA-EMA protocols. Biomedical LLMs showed lower relevance overall and frequently generated insufficient justification. All LLMs demonstrated limited proficiency in ontology-code mapping. CONCLUSION: Off-the-shelf general-purpose LLMs currently offered more reliable support for pharmacoepidemiologic study design than the smaller biomedical LLMs evaluated. In the model-controlled comparison (GPT-4o with Least-to-Most versus Active prompting), prompt strategy did not significantly affect performance (paired p = 0.93), and prompt comparisons are therefore reported descriptively. Because the two model groups also differed in scale and instruction-tuning, this contrast should be interpreted as a comparison of readily deployable options rather than of biomedical specialization alone.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Employing General-Purpose and Biomedical Large Language Models with Advanced Prompt Engineering for Pharmacoepidemiologic Study Design. — 科研速览 Science Skim