科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ American journal of health-system pharmacy : AJHP : official journal of the American Society of Health-System Pharmacists2026-09-10

Performance evaluation of a large language model for medication management tasks.

Kelli Henry, Steven Xu, Kaitlin Blotske, Moriah Cargile, Erin F Barreto, Brian Murray, Susan Smith, Seth R Bauer, Xingmeng Zhao, Adeleine Tilley, Yanjun Gao, Tianming Liu, Sunghwan Sohn, Andrea Sikora

一句话结论 · In one sentence

Model performance for basic medication tasks was consistently poor. This evaluation highlights the need for domain-specific training through clinician-annotated datasets and a comprehensive evaluation framework for benchmarking performance.

原始摘要(英文原文)· Original abstract
PURPOSE: Large language models (LLMs) have proven performance for certain diagnostic tasks; however, limited studies have evaluated their consistency in recommending appropriate medication regimens for a given diagnosis. Medication management is a complex task that requires synthesis of drug formulation and complete order instructions for safe use. Here, the performance of GPT-4o, an LLM available with OpenAI's ChatGPT, was tested on 3 medication management tasks. METHODS: GPT-4o performance was tested on 3 medication tasks: identifying available formulations for a given generic drug name, identifying drug-drug interactions (DDIs) for a given medication regimen, and preparing a medication order for a given generic drug name. For each experiment, the model's raw text response was captured exactly as returned and evaluated using clinician evaluation in addition to standard LLM metrics, including Term Frequency-Inverse Document Frequency (TF-IDF) vectors, normalized Levenshtein similarity, and Recall-Oriented Understudy for Gisting Evaluation (ROUGE-1/ROUGE-L) F1 score between each response and its reference string. RESULTS: For the first task of drug-formulation matching, GPT-4o had 49% accuracy for generic medications being matched to all available formulations, with an average of 1.23 omissions per medication and 1.14 hallucinations per medication. For the second task of drug-drug interaction identification, the accuracy was 54.7% for identifying the DDI pair. For the third task, GPT-4o generated order sentences containing no medication or abbreviation errors in 65.8% of the cases. CONCLUSION: Model performance for basic medication tasks was consistently poor. This evaluation highlights the need for domain-specific training through clinician-annotated datasets and a comprehensive evaluation framework for benchmarking performance.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Performance evaluation of a large language model for medication management tasks. — 科研速览 Science Skim