科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Military medicine2026-09-03

Evaluating the Risks, Safety, and Reliability of Large Language Models in Military Medicine.

Darshan Thota, David Alt, Sriram Venkatesan, Jesus Caban

一句话结论 · In one sentence

Failure to generate an accurate response can delay decision-making and increase cognitive burden and risk. Confidently incorrect responses are particularly dangerous. Although LLMs can generate useful clinical information, they remain vulnerable to biases and factual errors. Human oversight, training, and model tuning are necessary before utilizing military data.

原始摘要(英文原文)· Original abstract
OBJECTIVE: To evaluate the performance of Large Language Models (LLMs) within the Military Health System (MHS), specifically assessing for bias, hallucinations, and safety risks. METHODS: In late 2024, 47 clinicians across 13 military treatment facilities evaluated 8 blinded LLMs. Using standardized clinical vignettes, 835 conversations were reviewed for demographic bias, hallucinations, classification errors, and safety risks. RESULTS: Demographic bias was present in 82 conversations (9.8%), hallucinations in 85 (10.2%), classification errors in 19 (2.3%), and safety risks in 78 (9.3%). Classification errors were linked to military-specific jargon. Safety risks were most prominent in overdose, critical care, and psychiatric scenarios. CONCLUSION: Failure to generate an accurate response can delay decision-making and increase cognitive burden and risk. Confidently incorrect responses are particularly dangerous. Although LLMs can generate useful clinical information, they remain vulnerable to biases and factual errors. Human oversight, training, and model tuning are necessary before utilizing military data.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluating the Risks, Safety, and Reliability of Large Language Models in Military Medicine. — 科研速览 Science Skim